system
A multifunctional smartphone system with generative AI supports daily tasks like English conversation, cooking, and video editing by generating optimized content and feedback, addressing the inefficiencies of conventional systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2026-03-16
AI Technical Summary
Conventional smartphones lack efficient means to support various tasks in users' daily lives, particularly those requiring specialized knowledge such as English conversation practice, recipe suggestions, and video editing, with existing systems being inconvenient and ineffective.
A multifunctional smartphone system utilizing generative AI technology, comprising a terminal device for user input, a server for data generation, and an interface for user interaction, which generates optimized recipes, conversation scenarios, and video editing suggestions based on user requests.
Enables users to efficiently perform daily tasks like English conversation practice, cooking, and video editing by providing optimized content and feedback, enhancing user convenience and reducing effort.
Smart Images

Figure 2026047936000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional smartphones have lacked means to efficiently support various tasks in users' daily lives. In particular, for tasks that require specialized knowledge, such as English conversation practice, recipe suggestions, and video editing, there has been no convenient and effective way for users to perform them. Also, there is a need for a system to efficiently perform these tasks by effectively utilizing generative AI technology. Solving these problems and providing new value to users is an issue.
Means for Solving the Problems
[0005] To solve the above problems, the present invention provides the following means: a system including a terminal means for receiving requests from a user, a server means for receiving requests from the terminal means and providing data generated based on the requests, and a terminal means for displaying the generated data received from the server means and providing an interface for the user to perform the next operation. Furthermore, when the terminal means suggests a cooking recipe, it generates an optimal recipe based on the user's preferences and available ingredients, generates a conversation scenario using a generative AI to support English conversation practice, and provides feedback to the user.
[0006] "Terminal means" refers to a device used to receive requests from users or display data from a server, and includes, for example, smartphones and tablets.
[0007] A "server system" is a computer system that receives requests from terminal systems, processes the data generated based on those requests, and provides it.
[0008] A "request" is a request that a user sends to a server via a terminal device in order to perform a specific task or service.
[0009] "Generated data" refers to information and content created by the server using generation AI based on user requests, including, for example, English conversation scenarios, cooking recipes, and video editing suggestions.
[0010] An "interface" refers to the screen or operating means that a user uses to operate a terminal device, view generated data, and perform subsequent actions.
[0011] "Generative AI" refers to algorithms and systems that use artificial intelligence technology to automatically generate data and content based on user requests.
[0012] "English conversation practice" is an activity in which users use generative AI to simulate conversations and receive feedback in order to improve their English speaking and listening skills.
[0013] A "cooking recipe" is information that describes the steps and necessary ingredients for creating a specific dish, and is optimized based on the user's preferences and the ingredients they currently have on hand.
[0014] "Video editing" refers to the process of editing videos shot by users to improve them visually and audibly, using AI to provide editing suggestions and delivering the final edited result. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of the data processing device and smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention relates to a multifunctional smartphone system using generative AI, which enables users to efficiently perform various daily tasks. This system consists of "terminal means," "server means," and "generative AI technology."
[0037] First, the user operates the terminal device and requests a specific task (e.g., practicing English conversation, suggesting cooking recipes, video editing). For example, if the user wants to practice English conversation, they would use a dedicated app on the terminal device to input the request "Start practicing English conversation." Next, the terminal device sends this request to the server device. The server device receives the user's request and uses a generation AI to generate an appropriate conversation scenario. The generated conversation scenario is then sent back to the terminal device, and the user begins practicing English conversation accordingly.
[0038] Next, we will describe an example of a cooking recipe suggestion. When a user requests a cooking recipe, they input "a dish using tomatoes and chicken" using a dedicated app on their terminal device. The terminal device sends the request to the server device, which uses a generation AI to generate an optimal recipe based on the user's preferences and available ingredients. The generated recipe is then sent to the terminal device and displayed to the user. The user can then cook according to that recipe.
[0039] Next, we will explain a specific example of video editing. When a user wants to edit a video, they operate a terminal device to select the video they want to edit and request editing. The terminal device uploads the video material to a server device, and the server device uses AI generation to suggest editing options. For example, the server device might suggest "scene transition effects" or "adding music." The suggestions are sent to the terminal device, and the user reviews them. The user can then provide correction instructions as needed, and the final editing is completed by the server device. The completed video is sent back to the terminal device for the user to review.
[0040] To realize this system, the server is a computer system with high computing power and has software installed to execute generative AI technology. The terminal is equipped with an interface for user operation and has communication functions for communicating with the server.
[0041] The following is a specific example of a processing flow:
[0042] 1. The user operates a device and requests a specific task. (Examples: practicing English conversation, suggesting cooking recipes, video editing)
[0043] 2. The terminal sends the user's request to the server.
[0044] 3. The server receives the request and uses the generation AI to generate optimal data (e.g., conversation scenarios, cooking recipes, editing suggestions).
[0045] 4. The server transmits the generated data to the terminal device.
[0046] 5. The device displays the received data to the user and prompts them to take the next action.
[0047] 6. The user performs the following actions (e.g., responding to an English conversation, executing a cooking recipe, editing and correcting an edit).
[0048] Through the above process, a multi-functional smartphone system using generative AI can efficiently support the user's daily life.
[0049] The following describes the processing flow.
[0050] English conversation practice
[0051] Step 1:
[0052] The user wants to start practicing English conversation and operates the app on their device, clicking the "Practice English Conversation" button.
[0053] Step 2:
[0054] The terminal receives the user's request and sends the request to the server.
[0055] Step 3:
[0056] The server receives the request and activates the generation AI.
[0057] Step 4:
[0058] The server references user information (past English conversation history and level) and generates an appropriate conversation scenario.
[0059] Step 5:
[0060] The server generates a conversation scenario and sends it to the terminal device.
[0061] Step 6:
[0062] The device displays the conversation scenario it received to the user and presents the user with an initial question.
[0063] Step 7:
[0064] The user enters their answer to the question in text or voice.
[0065] Step 8:
[0066] The terminal sends the user's response to the server.
[0067] Step 9:
[0068] The server analyzes the user's responses and uses AI to generate appropriate next questions and feedback.
[0069] Step 10:
[0070] The server sends any new questions or feedback it generates to the terminal device.
[0071] Step 11:
[0072] The device displays new questions and feedback to the user.
[0073] Step 12:
[0074] The user responds again, and the process continues.
[0075] Cooking recipe suggestions
[0076] Step 1:
[0077] A user requests a cooking recipe and uses a dedicated app on their device to request a "tomato and chicken recipe."
[0078] Step 2:
[0079] The terminal receives the request and sends the request to the server.
[0080] Step 3:
[0081] The server receives the request and activates the generation AI.
[0082] Step 4:
[0083] The server collects information about the user's preferences and available ingredients to generate the optimal recipe.
[0084] Step 5:
[0085] The server sends the generated recipe to the terminal device.
[0086] Step 6:
[0087] The device displays the received recipe to the user.
[0088] Step 7:
[0089] The user views the displayed recipe and begins cooking.
[0090] Step 8:
[0091] The device displays instructions to the user for each step of the recipe.
[0092] Step 9:
[0093] The user follows the instructions and continues cooking.
[0094] Video editing
[0095] Step 1:
[0096] The user wants to edit a video and uses their device to select the video they want to edit.
[0097] Step 2:
[0098] The device uploads the selected video material to the server.
[0099] Step 3:
[0100] The server receives the video footage and activates the generation AI.
[0101] Step 4:
[0102] The server generates editing suggestions (e.g., scene transition effects or music additions) based on the video content.
[0103] Step 5:
[0104] The server sends the generated editing suggestions to the terminal device.
[0105] Step 6:
[0106] The device displays editing suggestions to the user.
[0107] Step 7:
[0108] The user reviews the proposal and indicates any necessary changes.
[0109] Step 8:
[0110] The server receives user instructions and performs the final editing.
[0111] Step 9:
[0112] The server renders the video and generates the final product.
[0113] Step 10:
[0114] The server sends the finished video to the terminal device.
[0115] Step 11:
[0116] The device displays the finished video to the user, who then reviews and saves it.
[0117] (Example 1)
[0118] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0119] Traditional smartphone applications lacked the automation of content generation based on user requests, resulting in significant user effort and time. Furthermore, they provided insufficient support for users to efficiently and effectively perform specific activities. In particular, traditional systems lacked flexibility and responsiveness for a wide range of tasks, such as practicing English conversation, getting cooking recipe suggestions, and video editing.
[0120] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0121] In this invention, the server includes a mobile terminal means for receiving requests from a user, a computer server means for receiving requests from the mobile terminal means and providing data generated based on the requests, a mobile terminal means for displaying the generated data received from the computer server means and providing an interface for the user to perform the next operation, a computer server means for generating optimal data based on the requests using a generation AI model, and a mobile terminal means for displaying specific instructions for the user to perform the next operation based on the generated data. This enables the user to efficiently perform a wide range of tasks, such as practicing English conversation, getting cooking recipe suggestions, and editing videos.
[0122] A "user" refers to a person who operates a mobile device to request a specific task.
[0123] "Mobile terminal means" refers to a portable electronic device that is operated by a user and has the function of communicating with a server.
[0124] A "computer server" refers to a high-performance computer system that has the function of processing user requests and providing generated data.
[0125] A "generative AI model" refers to artificial intelligence technology that generates optimal data based on user requests.
[0126] A "request" refers to a request for a specific task made by a user via a mobile device.
[0127] "Generated data" refers to information created by the computer server using a generation AI model, based on user requests.
[0128] "Interface" refers to the means of operation, such as screens and buttons, that users use to operate a mobile device.
[0129] Modes for carrying out the invention
[0130] This invention relates to a multifunctional smartphone system that enables users to efficiently perform a wide range of daily tasks. Specifically, it is a system that provides data generated based on requests from users, and is implemented using a mobile terminal, a computer server, and a generative AI model.
[0131] The system consists of the following main elements:
[0132] 1. Mobile device means:
[0133] The mobile device provides an interface for users to operate and enter requests. Users can enter various requests using a dedicated application. For example, requests such as "Start practicing English conversation," "Recipe a dish using tomatoes and chicken," or "Please give me suggestions for video editing" can be entered.
[0134] 2. Computer server means:
[0135] The computer server receives requests from users and generates optimal data using generative AI models. The server has high computing power and utilizes generative AI models such as OpenAI® GPT-3®.5. The generated data varies depending on the content of the request. For example, conversation scenarios are generated for English conversation practice, optimal recipes for cooking recipe suggestions, and editing suggestions for video editing.
[0136] 3. Generative AI Models:
[0137] The generative AI model is responsible for generating content based on user requests. It takes text-based prompts as input and generates relevant information. For example, it processes prompts such as "Suggest the best recipe using tomatoes and chicken" or "Generate a scenario suitable for practicing English conversation."
[0138] Specific examples are given below.
[0139] English conversation practice
[0140] The user launches a dedicated app on their smartphone and enters "Start practicing English conversation." The device sends this request to a server, which uses a generative AI model to generate an appropriate conversation scenario. The generated scenario is sent back to the device, and the user begins practicing English conversation according to that scenario.
[0141] Example prompt: "Start practicing English conversation."
[0142] Cooking recipe suggestions
[0143] The user enters "a dish using tomatoes and chicken" into a dedicated app. The device sends the request to the server, which uses a generative AI model to generate the optimal recipe based on the user's preferences and available ingredients. The generated recipe is sent back to the device, and the user cooks according to that recipe.
[0144] Example prompt: "A dish using tomatoes and chicken"
[0145] Video editing suggestions
[0146] The user selects the video they want to edit using a dedicated app and requests "video editing suggestions." The device uploads the video footage to a server, which uses a generation AI model to generate editing suggestions (e.g., scene transition effects, adding music). The generated editing suggestions are sent back to the device, where the user reviews them and provides correction instructions as needed.
[0147] Example of a prompt: "Please edit this video and add scene transition effects."
[0148] As described above, the coordinated operation of each element enables the optimal generation and provision of data in response to user requests. This invention is a system that efficiently and effectively supports the user's daily life.
[0149] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0150] Steps for practicing English conversation
[0151] Step 1:
[0152] The user enters a request to "start practicing English conversation."
[0153] Input: User request ("Start practicing English conversation")
[0154] Specific operation: The user inputs data using a dedicated app on their mobile device.
[0155] Step 2:
[0156] The terminal receives user input and sends a request to the server.
[0157] Input: User Request
[0158] Output: Request data from terminal to server
[0159] Specific operation: The terminal converts the request into the appropriate format and sends it to the server over the internet.
[0160] Step 3:
[0161] The server receives the request and prompts the generated AI model.
[0162] Input: Request received by the server
[0163] Output: Prompt message for the generating AI model ("Generate a scenario suitable for practicing English conversation")
[0164] Specific operation: The server analyzes the request content and creates and inputs a prompt sentence suitable for the generated AI model.
[0165] Step 4:
[0166] The generative AI model generates appropriate conversation scenarios.
[0167] Input: Prompt message
[0168] Output: Generated conversation scenario
[0169] Specific operation: The generative AI model (e.g., OpenAI GPT-3.5) generates conversation scenarios based on prompts.
[0170] Step 5:
[0171] The server sends the generated conversation scenario to the terminal.
[0172] Input: Generated conversation scenario
[0173] Output: Conversation scenario data from server to terminal
[0174] Specific operation: The server sends the generated scenario to the terminal in the appropriate format.
[0175] Step 6:
[0176] The device displays a conversation scenario to the user.
[0177] Input: Conversation scenario data from the server
[0178] Output: Conversation scenario displayed on the user's device
[0179] Specific operation: The terminal renders the received scenario on the display screen.
[0180] Step 7:
[0181] The user practices English conversation based on the displayed conversation scenario.
[0182] Input: Displayed conversation scenario
[0183] Output: User's English conversation response
[0184] Specific operation: The user responds through the terminal.
[0185] Processing steps for suggesting cooking recipes
[0186] Step 1:
[0187] The user enters a request for "a dish using tomatoes and chicken."
[0188] Input: User request ("A dish using tomatoes and chicken")
[0189] Specific operation: The user inputs data using a dedicated app on their mobile device.
[0190] Step 2:
[0191] The terminal receives user input and sends a request to the server.
[0192] Input: User Request
[0193] Output: Request data from terminal to server
[0194] Specific operation: The terminal converts the request into the appropriate format and sends it to the server over the internet.
[0195] Step 3:
[0196] The server receives the request and prompts the generated AI model.
[0197] Input: Request received by the server
[0198] Output: Prompt message for the generative AI model ("Suggest the best recipe using tomatoes and chicken as ingredients")
[0199] Specific operation: The server analyzes the request content and creates and inputs a prompt sentence suitable for the generated AI model.
[0200] Step 4:
[0201] The generative AI model generates the optimal cooking recipe.
[0202] Input: Prompt message
[0203] Output: Generated cooking recipe
[0204] Specific operation: The generative AI model generates cooking recipes based on prompts.
[0205] Step 5:
[0206] The server sends the generated cooking recipe to the terminal.
[0207] Input: Generated cooking recipe
[0208] Output: Recipe data from server to terminal
[0209] Specific operation: The server sends the generated recipe to the terminal in the appropriate format.
[0210] Step 6:
[0211] The device displays cooking recipes to the user.
[0212] Input: Recipe data from the server
[0213] Output: Recipe displayed on the user's device
[0214] Specific action: The device renders the received recipe on the display screen.
[0215] Step 7:
[0216] The user cooks according to the displayed recipe.
[0217] Input: Displayed recipe
[0218] Output: Cooked food
[0219] Specific actions: The user follows the displayed recipe and cooks the dish.
[0220] Processing steps for video editing suggestions
[0221] Step 1:
[0222] The user selects the video they want to edit and requests "video editing suggestions."
[0223] Input: User request ("I would like suggestions for video editing")
[0224] Specific operation: The user inputs data using a dedicated app on their mobile device.
[0225] Step 2:
[0226] The device uploads video footage to the server.
[0227] Input: User selected video footage
[0228] Output: Video data from terminal to server
[0229] Specific operation: The device compresses the video footage and sends it to the server via the internet.
[0230] Step 3:
[0231] The server receives the video footage and inputs prompts into the generating AI model.
[0232] Input: Video footage received by the server
[0233] Output: Prompt message for the generating AI model ("Generate editing suggestions suitable for the video")
[0234] Specific operation: The server analyzes the video material and creates and inputs prompt sentences suitable for the generating AI model.
[0235] Step 4:
[0236] The generative AI model generates appropriate editing suggestions.
[0237] Input: Prompt message
[0238] Output: Generated editing suggestions
[0239] Specific operation: The generation AI model generates editing suggestions (e.g., scene transition effects, adding music) based on prompts.
[0240] Step 5:
[0241] The server sends the generated editing suggestions to the terminal.
[0242] Input: Generated editing suggestions
[0243] Output: Editing suggestion data from server to terminal
[0244] Specific operation: The server sends the generated editing suggestions to the terminal in the appropriate format.
[0245] Step 6:
[0246] The device displays editing suggestions to the user.
[0247] Input: Editing suggestion data from the server
[0248] Output: Editing suggestions displayed on the user's device
[0249] Specific action: The terminal renders the received editing suggestions on the display screen.
[0250] Step 7:
[0251] The user reviews the displayed editing suggestions and provides correction instructions as needed.
[0252] Input: Displayed editing suggestions
[0253] Output: Correction instructions
[0254] Specific actions: The user reviews the displayed editing suggestions and enters correction instructions into the terminal as needed.
[0255] Step 8:
[0256] The server performs the final editing and sends the completed video to the device.
[0257] Input: Correction Instructions
[0258] Output: Completed video
[0259] Specific operation: The server performs the final editing based on the correction instructions and sends the completed video to the terminal.
[0260] Step 9:
[0261] The device displays the completed video to the user.
[0262] Input: Completed video data from the server
[0263] Output: The completed video will be displayed on the user's device.
[0264] Specific operation: The device renders the received completed video on the display screen and allows the user to confirm it.
[0265] (Application Example 1)
[0266] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0267] Modern consumers and store staff are required to efficiently handle a wide range of tasks, including providing real-time product information, suggesting optimal outfit combinations, and instantly announcing sales information. However, fulfilling these requirements with a single system is difficult and time-consuming, resulting in significant costs. This invention aims to improve convenience for both users and store staff by efficiently providing diverse information services required in physical stores using generative AI technology.
[0268] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0269] In this invention, the server includes terminal device means for receiving requests from users, server device means for receiving requests from the terminal device means and providing data generated based on the requests, and terminal device means for displaying the generated data received from the server device means and providing an interface for the user to perform subsequent operations. This makes it possible to provide product descriptions, coordination suggestions, and sales information.
[0270] A "user" is a person who uses a system to obtain information or request tasks.
[0271] A "terminal device means" is an electronic device used to receive requests from users or to display data from a server.
[0272] "Server device means" refers to a computer system that processes and provides data generated based on requests received from terminal device means.
[0273] "Generative AI technology" is artificial intelligence technology that generates optimal data and suggestions based on user requests.
[0274] "Product description" refers to providing users and store staff with detailed information about a specific product.
[0275] "Coordination suggestions" refer to providing optimal fashion and styling ideas by combining specific items.
[0276] "Sale information" refers to data about special prices and discounts currently available.
[0277] An "interface" is an operation screen or input / output means that allows a user to interact with a system and perform the following operations.
[0278] "Feedback" is the process of providing responses to user requests and related information.
[0279] The system for carrying out this invention comprises a terminal device means for receiving requests from users, a server device means for generating and providing data based on the requests, and a terminal device means for displaying the generated data to the user and providing an interface for performing subsequent operations. In this system, AI generation technology is utilized to generate optimal data in response to requests.
[0280] Explanation of the program's processing
[0281] 1. Terminal device means:
[0282] The terminal device means is an electronic device such as a smartphone or a tablet that receives requests from users. Using a dedicated application of the terminal device means, the user inputs a specific task (e.g., product description, coordination proposal, inquiry about sales information). For example, the user inputs a request such as "Please tell me the detailed product description of the XYZ smartwatch."
[0283] 2. Server device means:
[0284] The server device means receives the request sent from the front end and generates appropriate data using generative AI technology. The server system includes a computer system with high computing power, and software for executing the generative AI model (e.g., API of OpenAI) is installed. When the user's request is sent to the server, the server passes the prompt text to the generative AI and sends the obtained result back to the terminal device means.
[0285] As a specific example, when the user requests "Please propose a coordination using a red dress, black boots, and a white coat.", the server passes this prompt text to the generative AI and generates an appropriate coordination proposal.
[0286] 3. Interface providing means:
[0287] The terminal device means provides an interface for displaying the generated data received from the server to the user. Here, the user can check the generated data and perform the following operations. For example, the user can check the generated product description or coordination proposal, make a more detailed request if necessary, or consider purchasing at a physical store based on the provided information.
[0288] Prompt text of specific example
[0289] Product description:
[0290] "Could you please provide a detailed product description for the XYZ smartwatch?"
[0291] Outfit suggestions:
[0292] "Please suggest outfit ideas using a red dress, black boots, and a white coat."
[0293] This enables the provision of information that facilitates efficient communication between users and store staff in physical stores. Advanced information processing is achieved by utilizing generative AI models.
[0294] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0295] Step 1:
[0296] The user operates the terminal device and enters a specific request through a dedicated application. For example, if the user enters "Please tell me more about the XYZ smartwatch," this request is temporarily stored in the terminal device. The input data is saved as the request content and sent to the next processing step.
[0297] Step 2:
[0298] The terminal device transmits the received user request to the server device. The transmitted data is the user's request content (e.g., a request for product description), which is converted to an appropriate format and sent to the server. The input data is the user's request content, and the output data is the request data sent to the server.
[0299] Step 3:
[0300] The server device means analyzes the received request and generates appropriate data using generative AI technology. The generative AI model constructs a prompt sentence based on the request content and sends it to the generative AI. The input data is the user's request content and is processed as data into a prompt sentence. The output data obtained from the generative AI is specific information such as the generated product description.
[0301] Step 4:
[0302] The data obtained from the generative AI is processed by the server device means, converted into an appropriate format, and then sent to the terminal device means. The input data is the output data from the generative AI (e.g., product description), and the output data is the data converted into a format for transmission to the terminal device means.
[0303] Step 5:
[0304] The terminal device means displays the data received from the server to the user. The user checks the generated data (e.g., product description, coordination proposal, sales information). The input data is the generated data sent from the server, and the output data is the information displayed to the user.
[0305] Step 6:
[0306] The user checks the displayed information and performs the next operation. For example, when requesting more detailed information or making a decision to purchase a product at a physical store. The input data is the user's action based on the generated information, and the output data is to input the user's next operation into the terminal device means.
[0307] Through the above processing steps, a system is realized that efficiently provides data generated based on the user's request and improves the user's convenience.
[0308] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0309] This invention combines an emotion engine with a multifunctional smartphone system using generative AI, adjusting the system's responses and display content based on the user's emotional state. This system consists of a "terminal means," a "server means," "generative AI technology," and an "emotion engine."
[0310] First, the user operates the terminal device and requests a specific task (e.g., practicing English conversation, suggesting cooking recipes, video editing). For example, if the user wants to practice English conversation, they use a dedicated app on the terminal device to input the request "Practice English conversation." Next, the terminal device sends this request to the server device. The server device receives the user's request and activates the generative AI and emotion engine. The emotion engine recognizes the user's emotions from their facial expressions and tone of voice, and based on this information, the generative AI generates an appropriate conversation scenario. This scenario is then sent back to the terminal device, and the user begins practicing English conversation accordingly.
[0311] This document describes an example of a cooking recipe suggestion system. When a user requests a cooking recipe, they enter "a dish using tomatoes and chicken" into a dedicated app on their terminal device. The terminal device sends the request to a server device, which uses a generation AI and an emotion engine to generate an optimal recipe based on the user's preferences, available ingredients, and emotional state. The generated recipe is then sent to the terminal device and displayed to the user. The user can then cook according to the recipe. If the user is tired, the emotion engine can recognize this and suggest a simpler recipe.
[0312] Let's explain a specific example of video editing. When a user wants to edit a video, they operate a terminal device to select the video they want to edit and request editing. The terminal device uploads the video material to a server device, which uses a generation AI and emotion engine to suggest editing options. For example, the server device can analyze the user's emotional state and suggest simple editing options that minimize stress. The suggestions are sent to the terminal device, and the user reviews them. If necessary, the user provides correction instructions, and the final editing is completed by the server device. The completed video is sent back to the terminal device for the user to review.
[0313] The process would look like this:
[0314] First, the user operates a terminal device and requests a specific task (English conversation, recipe suggestions, video editing). The terminal device sends the request to the server device, which activates the generative AI and emotion engine. The emotion engine recognizes the user's emotional state and passes this information to the generative AI to generate the optimal scenario and suggestions. The generated data is sent to the terminal device and displayed to the user. The user then performs the next action based on this data. At each step, the emotion engine recognizes the user's emotional state in real time and adjusts the interface and response content accordingly.
[0315] To realize this system, the server must be a computer system with high computing power and have software installed to run the generative AI and emotion engine. Furthermore, the terminal must have an interface for user operation, communication capabilities for communicating with the server, and also be equipped with a camera and microphone for emotion recognition.
[0316] Thus, by combining a generative AI and an emotion engine, the system of the present invention can more precisely address the diverse needs of users, and as a result, improve the user experience.
[0317] The following describes the processing flow.
[0318] English conversation practice
[0319] Step 1:
[0320] A user wants to practice English conversation and clicks the "Practice English Conversation" button on the app on their device.
[0321] Step 2:
[0322] The terminal receives the user's request and sends the request to the server.
[0323] Step 3:
[0324] The server receives the request and activates the generative AI and emotion engine.
[0325] Step 4:
[0326] The emotion engine analyzes the user's facial expressions and voice tone transmitted from the device to recognize the user's emotional state.
[0327] Step 5:
[0328] The server uses AI to generate conversation scenarios based on user information (past English conversation history and level) and emotional state.
[0329] Step 6:
[0330] The server generates a conversation scenario and sends it to the terminal device.
[0331] Step 7:
[0332] The device displays the conversation scenario it received to the user and presents the user with an initial question.
[0333] Step 8:
[0334] The user enters their answer to the question in text or voice.
[0335] Step 9:
[0336] The terminal transmits the user's response, along with their facial expressions and tone of voice during the response, to the server.
[0337] Step 10:
[0338] The server analyzes the user's responses and emotional state, and then uses AI to generate appropriate questions and feedback.
[0339] Step 11:
[0340] The server sends any new questions or feedback it generates to the terminal device.
[0341] Step 12:
[0342] The device displays new questions and feedback to the user.
[0343] Step 13:
[0344] The user responds again, and the process continues.
[0345] Cooking recipe suggestions
[0346] Step 1:
[0347] A user requests a cooking recipe and uses a dedicated app on their device to request a "tomato and chicken recipe."
[0348] Step 2:
[0349] The terminal receives the request and sends the request to the server.
[0350] Step 3:
[0351] The server receives the request and activates the generative AI and emotion engine.
[0352] Step 4:
[0353] The emotion engine analyzes the user's facial expressions and voice tone transmitted from the device to recognize the user's emotional state.
[0354] Step 5:
[0355] The server uses AI to generate the optimal recipe based on the user's preferences, available ingredients, and emotional state.
[0356] Step 6:
[0357] The server sends the generated recipe to the terminal device.
[0358] Step 7:
[0359] The device displays the received recipe to the user. For example, it displays a "Tomato and Chicken Cream Stew Recipe".
[0360] Step 8:
[0361] The user checks the displayed recipe and begins cooking.
[0362] Step 9:
[0363] The device displays instructions to the user for each step of the recipe.
[0364] Step 10:
[0365] The user follows instructions and proceeds with cooking. Their facial expressions and tone of voice during cooking are also analyzed by an emotion engine.
[0366] Video editing
[0367] Step 1:
[0368] The user wants to edit a video and uses their device to select the video they want to edit.
[0369] Step 2:
[0370] The device uploads the selected video material to the server.
[0371] Step 3:
[0372] The server receives the video footage and activates the generation AI and emotion engine.
[0373] Step 4:
[0374] The emotion engine analyzes the user's facial expressions and voice transmitted from the device to recognize the user's emotional state.
[0375] Step 5:
[0376] The server uses AI to generate editing suggestions (e.g., scene transition effects or music additions) based on the video content and the user's emotional state.
[0377] Step 6:
[0378] The server sends the generated editing suggestions to the terminal device.
[0379] Step 7:
[0380] The device displays editing suggestions to the user. For example, it might show "recommended scene transition effects" or "add music."
[0381] Step 8:
[0382] The user reviews the suggestion and instructs on any necessary changes. For example, they might instruct to "change the music."
[0383] Step 9:
[0384] The server receives user instructions, and the final editing work is performed by a generation AI.
[0385] Step 10:
[0386] The server renders the video and generates the final product.
[0387] Step 11:
[0388] The server sends the finished video to the terminal device.
[0389] Step 12:
[0390] The device displays the finished video to the user, who then reviews, saves, or shares it.
[0391] (Example 2)
[0392] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0393] Traditional systems only provided standard responses and suggestions to user requests, making it difficult to offer personalized services that took into account the user's emotional state. This resulted in problems such as decreased user satisfaction and a lower quality of experience.
[0394] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0395] In this invention, the server includes information terminal device means for receiving requests from users, computer means for providing data generated based on the requests, information terminal device means for displaying the generated data and providing an interface for the user to perform the next operation, emotion analysis means for recognizing the user's emotional state, and generation AI means for generating conversation scenarios and suggestions based on the emotional state. This makes it possible to provide highly personalized services that respond to the user's emotional state.
[0396] An "information terminal device means" is a device that a user operates and inputs requests into, and that is equipped with communication functions and a user interface for sending requests.
[0397] A "computer means" is a high-performance computer system that receives requests, processes data using generative AI technology and sentiment analysis technology, and provides the generated data.
[0398] "Emotional analysis means" refers to technologies and devices that analyze a user's facial expressions, tone of voice, and other emotional indicators to recognize the user's emotional state.
[0399] "Generative AI means" refers to technologies and devices that use generative AI models to generate scenarios and suggestions based on user requests.
[0400] An "interface" refers to a system that includes user interfaces and input devices for reviewing generated data and performing subsequent actions.
[0401] The present invention is a system that provides services and responses based on the user's emotional state. This system is composed of an information terminal device, a computer, an emotion analysis device, a generation AI device, and an interface.
[0402] When a user requests a specific task using the information terminal device, the information terminal device transmits this request to the computer device. The computer device processes the received request and activates the generation AI device and the emotion analysis device. The emotion analysis device analyzes emotional indicators such as the user's facial expressions and tone of voice to recognize the user's emotional state. Subsequently, the generation AI device generates optimal scenarios and suggestions based on the recognized emotional state. The generated scenarios and suggestions are then transmitted back to the information terminal device device by the computer device and displayed to the user.
[0403] The following specific hardware and software can be used in this invention:
[0404] Information terminal device means: For example, smartphones (iPhone® and Android® devices) are equipped with dedicated apps and have built-in cameras and microphones.
[0405] Computing equipment: High-performance computer systems (e.g., AWS®, Google® Cloud) are installed with software for running generative AI models and sentiment analysis techniques.
[0406] Emotion analysis methods include facial recognition technology (e.g., Microsoft® Azure® facial recognition API) and speech analysis technology (e.g., Google Cloud Speech-to-Text API).
[0407] Generative AI method: Use a generative AI model (e.g., GPT series) to generate scenarios and suggestions based on user requests.
[0408] Specific example
[0409] English conversation practice
[0410] 1. The user launches the dedicated app on their smartphone and enters "I want to practice English conversation."
[0411] 2. The information terminal device transmits this request to the computer.
[0412] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[0413] 4. The emotion analysis system analyzes the user's facial expressions and voice to recognize that they are relaxed.
[0414] 5. The generation AI generates a relaxed conversation scenario and sends it to the computer.
[0415] 6. The computer means transmits the generated scenario to the information terminal device means.
[0416] 7. The information terminal device displays the conversation scenario to the user.
[0417] 8. The user practices English conversation following the provided scenario.
[0418] Cooking recipe suggestions
[0419] 1. The user enters "I would like suggestions for dishes using tomatoes and chicken."
[0420] 2. The information terminal device transmits a request to the computer.
[0421] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[0422] 4. The emotion analysis tool analyzes the user's emotional state and, for example, recognizes that the user is slightly tired.
[0423] 5. The AI generates simple cooking recipes that require little physical effort and sends them to the computer.
[0424] 6. The computer means transmits the generated recipe to the information terminal device means.
[0425] 7. The information terminal device displays the recipe to the user.
[0426] 8. The user begins cooking according to the suggested recipe.
[0427] Video editing
[0428] 1. The user selects the video they want to edit and requests, "Easily edit this video."
[0429] 2. The information terminal device uploads the video to the computer and sends an editing request.
[0430] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[0431] 4. The emotion analysis tool analyzes the user's emotional state and recognizes that they are experiencing stress.
[0432] 5. The generation AI means generates simple editing suggestions and sends them to the computer means.
[0433] 6. The computer means transmits the generated editing proposal to the information terminal device means.
[0434] 7. The information terminal device displays editing suggestions to the user.
[0435] 8. Review the editing suggestions provided by the user and make any necessary corrections.
[0436] Examples of prompt statements include:
[0437] English conversation practice: "I want to practice speaking English."
[0438] Recipe suggestion: "Please suggest a dish using tomatoes and chicken."
[0439] Video editing request: "Please edit this video easily."
[0440] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0441] Step 1:
[0442] The user operates an information terminal device and requests a specific task.
[0443] Input: The user launches the dedicated app and enters a prompt message (e.g., "I want to practice English conversation").
[0444] Output: The information terminal device retrieves the user's request data.
[0445] Specific action: The user submits a request using an input form on their smartphone screen.
[0446] Step 2:
[0447] The terminal sends a request to the server.
[0448] Input: Retrieved request data.
[0449] Output: Sends the request data to the server.
[0450] Specific operation: The information terminal device sends request data to the server via an HTTP request. For example, it POSTs the request data to the endpoint.
[0451] Step 3:
[0452] The server processes the request and activates the generative AI and emotion analysis tools.
[0453] Input: Request data sent from the terminal.
[0454] Output: Requirements for emotion analysis and generation AI means.
[0455] Specific operation: The server analyzes the request data and selects the appropriate processing pipeline. It initializes and starts the generative AI and sentiment analysis tools.
[0456] Step 4:
[0457] The emotion analysis tool recognizes the user's emotional state.
[0458] Input: User's facial expression images and audio data sent to the server.
[0459] Output: Data indicating emotional state (e.g., relaxed, tense, tired).
[0460] Specific operation: The emotion analysis system analyzes the user's facial expressions and voice, and evaluates their emotional state using facial recognition technology and voice tone analysis technology.
[0461] Step 5:
[0462] The generation AI generates appropriate scenarios and suggestions.
[0463] Input: Request data, emotion state data.
[0464] Output: Generated scenarios and suggested data.
[0465] Specific operation: The generative AI uses models such as the GPT series to create optimal scenarios and suggestions by combining user requests and emotional states.
[0466] Step 6:
[0467] The server sends the generated scenarios and suggestions to the terminal.
[0468] Input: Scenarios and suggested data generated by the AI.
[0469] Output: Scenario and suggestion data sent to the terminal.
[0470] Specific operation: The server sends the generated scenarios and suggestions to the terminal as an HTTP response.
[0471] Step 7:
[0472] The device displays scenarios and suggestions it has received to the user.
[0473] Input: Scenario and suggestion data sent from the server.
[0474] Output: User screen displaying scenarios and suggestions.
[0475] Specific operation: The device analyzes data and displays it on the user interface. For example, it displays a conversation scenario in a chat format.
[0476] Step 8:
[0477] Users perform tasks according to scenarios and suggestions.
[0478] Input: Scenario and suggested data displayed on the terminal.
[0479] Output: User task execution status.
[0480] Specific actions: The user follows the presented scenarios and suggestions to perform the specified tasks (e.g., practicing English conversation, cooking).
[0481] Step 9:
[0482] At each step, the emotion analysis system recognizes the user's emotional state in real time and adjusts the interface and responses accordingly.
[0483] Input: User's current emotional state data.
[0484] Output: Updates to the user interface and response content.
[0485] Specific operation: The emotion analysis system monitors the emotional state in real time and adjusts the displayed content and responses as needed. For example, if it detects that the user is tired, it displays a message prompting a simple action.
[0486] By clearly specifying the concrete inputs, outputs, and actions at each step, the system's processing logic can be explained in detail.
[0487] (Application Example 2)
[0488] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0489] Traditional content delivery services lacked a mechanism to provide optimal content based on the user's emotional state, making it difficult to appropriately recommend content that users were looking for at that moment. This could lead to a lack of improvement in the user experience and a decrease in user satisfaction.
[0490] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0491] In this invention, the server includes terminal means for receiving requests from users, server means for receiving requests from the terminal means and providing data generated based on the requests, terminal means for displaying the generated data received from the server means and providing an interface for the user to perform the next operation, an emotion engine for analyzing the user's facial image and voice data and recognizing their emotional state, and generative AI technology for generating optimal content recommendations based on the output of the emotion engine. This makes it possible to recommend optimal content according to the user's emotional state.
[0492] A "terminal device" is a device used to receive requests from users.
[0493] A "server device" is a device that provides data generated based on a request received from a terminal device.
[0494] An "interface" is a mechanism that displays the generated data received from a server and allows the user to perform the following operations.
[0495] An "emotion engine" is a system that analyzes a user's facial image and voice data to recognize their emotional state.
[0496] "Generative AI technology" is an artificial intelligence technology that generates optimal content recommendations based on the output of an emotion engine.
[0497] "Content recommendation" refers to the suggestion of optimal content provided while taking the user's emotional state into consideration.
[0498] This invention is a system that combines generative AI technology and an emotion engine to recommend optimal content based on the user's emotional state. The system consists of multiple terminal means, server means, an emotion engine, and generative AI technology.
[0499] System program
[0500] The system basically operates in the following way:
[0501] 1. A device (such as a smartphone or tablet) receives requests from the user. This includes actions by the user to request content (such as playing a video or entering a search query).
[0502] 2. The terminal device acquires the user's facial image and voice, and collects data for emotion recognition. A digital camera and microphone are used in this process.
[0503] 3. The server receives requests and emotion recognition data from the terminal. It then analyzes the user's emotional state using an emotion engine. The emotion engine uses facial recognition software such as "DeepFace".
[0504] 4. The server uses generative AI technology based on the emotional state and request content to recommend the most suitable content. This process utilizes generative AI models such as the language model "GPT-3". For example, if a user is in a "happy mood" and is looking for recommended videos, this generative AI model will generate appropriate prompts and suggest content.
[0505] Hardware and software
[0506] Hardware:
[0507] The device must be equipped with a camera and a microphone. This allows for real-time analysis of the user's facial expressions and voice.
[0508] software:
[0509] The emotion engine uses facial recognition software such as "DeepFace." The generative AI technology uses large-scale language models such as "GPT-3."
[0510] Examples of specific cases and prompt statements
[0511] Let's consider what makes a user happy and what kind of video they would enjoy watching.
[0512] Specific example:
[0513] The terminal device captures the user's facial image, and the emotion engine recognizes the emotion "happy." This information is sent to the server, and the generating AI model recommends the most suitable content based on the following prompt message.
[0514] Example of a prompt:
[0515] "Please recommend videos that best capture the 'happiness' that users are experiencing."
[0516] This recommendation system enables personalized content delivery based on the user's emotional state, resulting in a more comfortable user experience.
[0517] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0518] Step 1:
[0519] The user operates a device such as a smartphone to enter a request (e.g., video recommendations tailored to their emotions). The request data is entered into the device and sent to the next step.
[0520] Step 2:
[0521] The terminal device captures the user's facial image and voice data using a camera and microphone. The captured facial image and voice data become input data for analysis by the emotion engine. Digital image processing and voice analysis are performed here.
[0522] Step 3:
[0523] The terminal device sends request data and facial image / audio data to the server device. The server device receives this data and performs emotion analysis using an emotion engine. The input data is the request and emotion analysis data, and the output data is the user's emotional state (e.g., "happy").
[0524] Step 4:
[0525] The server uses an emotion engine to analyze facial images and audio data to identify the user's emotional state. Specifically, it uses facial recognition software such as "DeepFace" to determine the emotional state. The data processing performed in this process involves emotion estimation through image and audio analysis.
[0526] Step 5:
[0527] Based on the emotion engine's output data (the user's emotional state), the server uses generative AI technology to generate optimal content recommendations. Specifically, it uses the language model "GPT-3" to generate prompt sentences that correspond to the user's emotions and then makes content recommendations based on those prompts. This involves both prompt sentence generation and the generation of recommendation data using a large-scale language model.
[0528] Step 6:
[0529] The server sends content recommendation data generated by generative AI technology to the terminal device. The terminal device receives the recommendation data and displays it to the user through its interface. The data processing performed here is UI rendering, which takes the recommendation data from the server and displays it in an easy-to-understand manner for the user.
[0530] Step 7:
[0531] The user reviews the displayed content recommendations and selects the next action (e.g., playing a video). This provides feedback that improves the user experience and leads to further requests. In this step, a loop of data collection and feedback based on user interaction is crucial.
[0532] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0533] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0534] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0535] [Second Embodiment]
[0536] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0537] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0538] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0539] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0540] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0541] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0542] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0543] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0544] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0545] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0546] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0547] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0548] This invention relates to a multifunctional smartphone system using generative AI, which enables users to efficiently perform various daily tasks. This system consists of "terminal means," "server means," and "generative AI technology."
[0549] First, the user operates the terminal device and requests a specific task (e.g., practicing English conversation, suggesting cooking recipes, video editing). For example, if the user wants to practice English conversation, they would use a dedicated app on the terminal device to input the request "Start practicing English conversation." Next, the terminal device sends this request to the server device. The server device receives the user's request and uses a generation AI to generate an appropriate conversation scenario. The generated conversation scenario is then sent back to the terminal device, and the user begins practicing English conversation accordingly.
[0550] Next, we will describe an example of a cooking recipe suggestion. When a user requests a cooking recipe, they input "a dish using tomatoes and chicken" using a dedicated app on their terminal device. The terminal device sends the request to the server device, which uses a generation AI to generate an optimal recipe based on the user's preferences and available ingredients. The generated recipe is then sent to the terminal device and displayed to the user. The user can then cook according to that recipe.
[0551] Next, we will explain a specific example of video editing. When a user wants to edit a video, they operate a terminal device to select the video they want to edit and request editing. The terminal device uploads the video material to a server device, and the server device uses AI generation to suggest editing options. For example, the server device might suggest "scene transition effects" or "adding music." The suggestions are sent to the terminal device, and the user reviews them. The user can then provide correction instructions as needed, and the final editing is completed by the server device. The completed video is sent back to the terminal device for the user to review.
[0552] To realize this system, the server is a computer system with high computing power and has software installed to execute generative AI technology. The terminal is equipped with an interface for user operation and has communication functions for communicating with the server.
[0553] The following is a specific example of a processing flow:
[0554] 1. The user operates a device and requests a specific task. (Examples: practicing English conversation, suggesting cooking recipes, video editing)
[0555] 2. The terminal sends the user's request to the server.
[0556] 3. The server receives the request and uses the generation AI to generate optimal data (e.g., conversation scenarios, cooking recipes, editing suggestions).
[0557] 4. The server transmits the generated data to the terminal device.
[0558] 5. The device displays the received data to the user and prompts them to take the next action.
[0559] 6. The user performs the following actions (e.g., responding to an English conversation, executing a cooking recipe, editing and correcting an edit).
[0560] Through the above process, a multi-functional smartphone system using generative AI can efficiently support the user's daily life.
[0561] The following describes the processing flow.
[0562] English conversation practice
[0563] Step 1:
[0564] The user wants to start practicing English conversation and operates the app on their device, clicking the "Practice English Conversation" button.
[0565] Step 2:
[0566] The terminal receives the user's request and sends the request to the server.
[0567] Step 3:
[0568] The server receives the request and activates the generation AI.
[0569] Step 4:
[0570] The server references user information (past English conversation history and level) and generates an appropriate conversation scenario.
[0571] Step 5:
[0572] The server generates a conversation scenario and sends it to the terminal device.
[0573] Step 6:
[0574] The device displays the conversation scenario it received to the user and presents the user with an initial question.
[0575] Step 7:
[0576] The user enters their answer to the question in text or voice.
[0577] Step 8:
[0578] The terminal sends the user's response to the server.
[0579] Step 9:
[0580] The server analyzes the user's responses and uses AI to generate appropriate next questions and feedback.
[0581] Step 10:
[0582] The server sends any new questions or feedback it generates to the terminal device.
[0583] Step 11:
[0584] The device displays new questions and feedback to the user.
[0585] Step 12:
[0586] The user responds again, and the process continues.
[0587] Cooking recipe suggestions
[0588] Step 1:
[0589] A user requests a cooking recipe and uses a dedicated app on their device to request a "tomato and chicken recipe."
[0590] Step 2:
[0591] The terminal receives the request and sends the request to the server.
[0592] Step 3:
[0593] The server receives the request and activates the generation AI.
[0594] Step 4:
[0595] The server collects information about the user's preferences and available ingredients to generate the optimal recipe.
[0596] Step 5:
[0597] The server sends the generated recipe to the terminal device.
[0598] Step 6:
[0599] The device displays the received recipe to the user.
[0600] Step 7:
[0601] The user views the displayed recipe and begins cooking.
[0602] Step 8:
[0603] The device displays instructions to the user for each step of the recipe.
[0604] Step 9:
[0605] The user follows the instructions and continues cooking.
[0606] Video editing
[0607] Step 1:
[0608] The user wants to edit a video and uses their device to select the video they want to edit.
[0609] Step 2:
[0610] The device uploads the selected video material to the server.
[0611] Step 3:
[0612] The server receives the video footage and activates the generation AI.
[0613] Step 4:
[0614] The server generates editing suggestions (e.g., scene transition effects or music additions) based on the video content.
[0615] Step 5:
[0616] The server sends the generated editing suggestions to the terminal device.
[0617] Step 6:
[0618] The device displays editing suggestions to the user.
[0619] Step 7:
[0620] The user reviews the proposal and indicates any necessary changes.
[0621] Step 8:
[0622] The server receives user instructions and performs the final editing.
[0623] Step 9:
[0624] The server renders the video and generates the final product.
[0625] Step 10:
[0626] The server sends the finished video to the terminal device.
[0627] Step 11:
[0628] The device displays the finished video to the user, who then reviews and saves it.
[0629] (Example 1)
[0630] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0631] Traditional smartphone applications lacked the automation of content generation based on user requests, resulting in significant user effort and time. Furthermore, they provided insufficient support for users to efficiently and effectively perform specific activities. In particular, traditional systems lacked flexibility and responsiveness for a wide range of tasks, such as practicing English conversation, getting cooking recipe suggestions, and video editing.
[0632] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0633] In this invention, the server includes a mobile terminal means for receiving requests from a user, a computer server means for receiving requests from the mobile terminal means and providing data generated based on the requests, a mobile terminal means for displaying the generated data received from the computer server means and providing an interface for the user to perform the next operation, a computer server means for generating optimal data based on the requests using a generation AI model, and a mobile terminal means for displaying specific instructions for the user to perform the next operation based on the generated data. This enables the user to efficiently perform a wide range of tasks, such as practicing English conversation, getting cooking recipe suggestions, and editing videos.
[0634] A "user" refers to a person who operates a mobile device to request a specific task.
[0635] "Mobile terminal means" refers to a portable electronic device that is operated by a user and has the function of communicating with a server.
[0636] A "computer server" refers to a high-performance computer system that has the function of processing user requests and providing generated data.
[0637] A "generative AI model" refers to artificial intelligence technology that generates optimal data based on user requests.
[0638] A "request" refers to a request for a specific task made by a user via a mobile device.
[0639] "Generated data" refers to information created by the computer server using a generation AI model, based on user requests.
[0640] "Interface" refers to the means of operation, such as screens and buttons, that users use to operate a mobile device.
[0641] Modes for carrying out the invention
[0642] This invention relates to a multifunctional smartphone system that enables users to efficiently perform a wide range of daily tasks. Specifically, it is a system that provides data generated based on requests from users, and is implemented using a mobile terminal, a computer server, and a generative AI model.
[0643] The system consists of the following main elements:
[0644] 1. Mobile device means:
[0645] The mobile device provides an interface for users to operate and enter requests. Users can enter various requests using a dedicated application. For example, requests such as "Start practicing English conversation," "Recipe a dish using tomatoes and chicken," or "Please give me suggestions for video editing" can be entered.
[0646] 2. Computer server means:
[0647] The computing server receives requests from users and generates optimal data using a generative AI model. The server has high computing power and utilizes generative AI models such as OpenAI GPT-3.5. The generated data varies depending on the content of the request. For example, conversation scenarios are generated for English conversation practice, optimal recipes for cooking recipe suggestions, and editing suggestions for video editing.
[0648] 3. Generative AI Models:
[0649] The generative AI model is responsible for generating content based on user requests. It takes text-based prompts as input and generates relevant information. For example, it processes prompts such as "Suggest the best recipe using tomatoes and chicken" or "Generate a scenario suitable for practicing English conversation."
[0650] Specific examples are given below.
[0651] English conversation practice
[0652] The user launches a dedicated app on their smartphone and enters "Start practicing English conversation." The device sends this request to a server, which uses a generative AI model to generate an appropriate conversation scenario. The generated scenario is sent back to the device, and the user begins practicing English conversation according to that scenario.
[0653] Example prompt: "Start practicing English conversation."
[0654] Cooking recipe suggestions
[0655] The user enters "a dish using tomatoes and chicken" into a dedicated app. The device sends the request to the server, which uses a generative AI model to generate the optimal recipe based on the user's preferences and available ingredients. The generated recipe is sent back to the device, and the user cooks according to that recipe.
[0656] Example prompt: "A dish using tomatoes and chicken"
[0657] Video editing suggestions
[0658] The user selects the video they want to edit using a dedicated app and requests "video editing suggestions." The device uploads the video footage to a server, which uses a generation AI model to generate editing suggestions (e.g., scene transition effects, adding music). The generated editing suggestions are sent back to the device, where the user reviews them and provides correction instructions as needed.
[0659] Example of a prompt: "Please edit this video and add scene transition effects."
[0660] As described above, the coordinated operation of each element enables the optimal generation and provision of data in response to user requests. This invention is a system that efficiently and effectively supports the user's daily life.
[0661] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0662] Steps for practicing English conversation
[0663] Step 1:
[0664] The user enters a request to "start practicing English conversation."
[0665] Input: User request ("Start practicing English conversation")
[0666] Specific operation: The user inputs data using a dedicated app on their mobile device.
[0667] Step 2:
[0668] The terminal receives user input and sends a request to the server.
[0669] Input: User Request
[0670] Output: Request data from terminal to server
[0671] Specific operation: The terminal converts the request into the appropriate format and sends it to the server over the internet.
[0672] Step 3:
[0673] The server receives the request and prompts the generated AI model.
[0674] Input: Request received by the server
[0675] Output: Prompt message for the generating AI model ("Generate a scenario suitable for practicing English conversation")
[0676] Specific operation: The server analyzes the request content and creates and inputs a prompt sentence suitable for the generated AI model.
[0677] Step 4:
[0678] The generative AI model generates appropriate conversation scenarios.
[0679] Input: Prompt message
[0680] Output: Generated conversation scenario
[0681] Specific operation: The generative AI model (e.g., OpenAI GPT-3.5) generates conversation scenarios based on prompts.
[0682] Step 5:
[0683] The server sends the generated conversation scenario to the terminal.
[0684] Input: Generated conversation scenario
[0685] Output: Conversation scenario data from server to terminal
[0686] Specific operation: The server sends the generated scenario to the terminal in the appropriate format.
[0687] Step 6:
[0688] The device displays a conversation scenario to the user.
[0689] Input: Conversation scenario data from the server
[0690] Output: Conversation scenario displayed on the user's device
[0691] Specific operation: The terminal renders the received scenario on the display screen.
[0692] Step 7:
[0693] The user practices English conversation based on the displayed conversation scenario.
[0694] Input: Displayed conversation scenario
[0695] Output: User's English conversation response
[0696] Specific operation: The user responds through the terminal.
[0697] Processing steps for suggesting cooking recipes
[0698] Step 1:
[0699] The user enters a request for "a dish using tomatoes and chicken."
[0700] Input: User request ("A dish using tomatoes and chicken")
[0701] Specific operation: The user inputs data using a dedicated app on their mobile device.
[0702] Step 2:
[0703] The terminal receives user input and sends a request to the server.
[0704] Input: User Request
[0705] Output: Request data from terminal to server
[0706] Specific operation: The terminal converts the request into the appropriate format and sends it to the server over the internet.
[0707] Step 3:
[0708] The server receives the request and prompts the generated AI model.
[0709] Input: Request received by the server
[0710] Output: Prompt message for the generative AI model ("Suggest the best recipe using tomatoes and chicken as ingredients")
[0711] Specific operation: The server analyzes the request content and creates and inputs a prompt sentence suitable for the generated AI model.
[0712] Step 4:
[0713] The generative AI model generates the optimal cooking recipe.
[0714] Input: Prompt message
[0715] Output: Generated cooking recipe
[0716] Specific operation: The generative AI model generates cooking recipes based on prompts.
[0717] Step 5:
[0718] The server sends the generated cooking recipe to the terminal.
[0719] Input: Generated cooking recipe
[0720] Output: Recipe data from server to terminal
[0721] Specific operation: The server sends the generated recipe to the terminal in the appropriate format.
[0722] Step 6:
[0723] The device displays cooking recipes to the user.
[0724] Input: Recipe data from the server
[0725] Output: Recipe displayed on the user's device
[0726] Specific action: The device renders the received recipe on the display screen.
[0727] Step 7:
[0728] The user cooks according to the displayed recipe.
[0729] Input: Displayed recipe
[0730] Output: Cooked food
[0731] Specific actions: The user follows the displayed recipe and cooks the dish.
[0732] Processing steps for video editing suggestions
[0733] Step 1:
[0734] The user selects the video they want to edit and requests "video editing suggestions."
[0735] Input: User request ("I would like suggestions for video editing")
[0736] Specific operation: The user inputs data using a dedicated app on their mobile device.
[0737] Step 2:
[0738] The device uploads video footage to the server.
[0739] Input: User selected video footage
[0740] Output: Video data from terminal to server
[0741] Specific operation: The device compresses the video footage and sends it to the server via the internet.
[0742] Step 3:
[0743] The server receives the video footage and inputs prompts into the generating AI model.
[0744] Input: Video footage received by the server
[0745] Output: Prompt message for the generating AI model ("Generate editing suggestions suitable for the video")
[0746] Specific operation: The server analyzes the video material and creates and inputs prompt sentences suitable for the generating AI model.
[0747] Step 4:
[0748] The generative AI model generates appropriate editing suggestions.
[0749] Input: Prompt message
[0750] Output: Generated editing suggestions
[0751] Specific operation: The generation AI model generates editing suggestions (e.g., scene transition effects, adding music) based on prompts.
[0752] Step 5:
[0753] The server sends the generated editing suggestions to the terminal.
[0754] Input: Generated editing suggestions
[0755] Output: Editing suggestion data from server to terminal
[0756] Specific operation: The server sends the generated editing suggestions to the terminal in the appropriate format.
[0757] Step 6:
[0758] The device displays editing suggestions to the user.
[0759] Input: Editing suggestion data from the server
[0760] Output: Editing suggestions displayed on the user's device
[0761] Specific action: The terminal renders the received editing suggestions on the display screen.
[0762] Step 7:
[0763] The user reviews the displayed editing suggestions and provides correction instructions as needed.
[0764] Input: Displayed editing suggestions
[0765] Output: Correction instructions
[0766] Specific actions: The user reviews the displayed editing suggestions and enters correction instructions into the terminal as needed.
[0767] Step 8:
[0768] The server performs the final editing and sends the completed video to the device.
[0769] Input: Correction Instructions
[0770] Output: Completed video
[0771] Specific operation: The server performs the final editing based on the correction instructions and sends the completed video to the terminal.
[0772] Step 9:
[0773] The device displays the completed video to the user.
[0774] Input: Completed video data from the server
[0775] Output: The completed video will be displayed on the user's device.
[0776] Specific operation: The device renders the received completed video on the display screen and allows the user to confirm it.
[0777] (Application Example 1)
[0778] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0779] Modern consumers and store staff are required to efficiently handle a wide range of tasks, including providing real-time product information, suggesting optimal outfit combinations, and instantly announcing sales information. However, fulfilling these requirements with a single system is difficult and time-consuming, resulting in significant costs. This invention aims to improve convenience for both users and store staff by efficiently providing diverse information services required in physical stores using generative AI technology.
[0780] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0781] In this invention, the server includes terminal device means for receiving requests from users, server device means for receiving requests from the terminal device means and providing data generated based on the requests, and terminal device means for displaying the generated data received from the server device means and providing an interface for the user to perform subsequent operations. This makes it possible to provide product descriptions, coordination suggestions, and sales information.
[0782] A "user" is a person who uses a system to obtain information or request tasks.
[0783] A "terminal device means" is an electronic device used to receive requests from users or to display data from a server.
[0784] "Server device means" refers to a computer system that processes and provides data generated based on requests received from terminal device means.
[0785] "Generative AI technology" is artificial intelligence technology that generates optimal data and suggestions based on user requests.
[0786] "Product description" refers to providing users and store staff with detailed information about a specific product.
[0787] "Coordination suggestions" refer to providing optimal fashion and styling ideas by combining specific items.
[0788] "Sale information" refers to data about special prices and discounts currently available.
[0789] An "interface" is an operation screen or input / output means that allows a user to interact with a system and perform the following operations.
[0790] "Feedback" is the process of providing responses to user requests and related information.
[0791] The system for carrying out this invention comprises a terminal device means for receiving requests from users, a server device means for generating and providing data based on the requests, and a terminal device means for displaying the generated data to the user and providing an interface for performing subsequent operations. In this system, AI generation technology is utilized to generate optimal data in response to requests.
[0792] Explanation of the program's processing
[0793] 1. Terminal device means:
[0794] The terminal device is an electronic device such as a smartphone or tablet that receives requests from the user. The user uses a dedicated application on the terminal device to input specific tasks (e.g., product description, outfit suggestions, sales information inquiries). For example, the user inputs a request such as, "Please tell me more about the XYZ smartwatch."
[0795] 2. Server device means:
[0796] The server device receives requests sent from the frontend and generates appropriate data using generative AI technology. The server system includes a computer system with high computing power and has software installed to run generative AI models (e.g., OpenAI API). When a user request is sent to the server, the server passes a prompt message to the generative AI and sends the obtained result back to the terminal device.
[0797] For example, if a user requests, "Please suggest an outfit using a red dress, black boots, and a white coat," the server will pass this prompt to the AI that generates the appropriate outfit suggestions.
[0798] 3. Means of providing the interface:
[0799] The terminal device provides an interface that displays the generated data received from the server to the user. The user can then review the generated data and perform the following actions: for example, review the generated product descriptions and styling suggestions, make more detailed requests as needed, or consider purchasing items at a physical store based on the provided information.
[0800] Example prompt statements
[0801] Product Description:
[0802] "Could you please provide a detailed product description for the XYZ smartwatch?"
[0803] Outfit suggestions:
[0804] "Please suggest outfit ideas using a red dress, black boots, and a white coat."
[0805] This enables the provision of information that facilitates efficient communication between users and store staff in physical stores. Advanced information processing is achieved by utilizing generative AI models.
[0806] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0807] Step 1:
[0808] The user operates the terminal device and enters a specific request through a dedicated application. For example, if the user enters "Please tell me more about the XYZ smartwatch," this request is temporarily stored in the terminal device. The input data is saved as the request content and sent to the next processing step.
[0809] Step 2:
[0810] The terminal device transmits the received user request to the server device. The transmitted data is the user's request content (e.g., a request for product description), which is converted to an appropriate format and sent to the server. The input data is the user's request content, and the output data is the request data sent to the server.
[0811] Step 3:
[0812] The server device analyzes the received request and generates appropriate data using generative AI technology. The generative AI model constructs a prompt sentence based on the request content and sends it to the generative AI. The input data is the user's request content, which is processed into a prompt sentence. The output data obtained from the generative AI is specific information such as the generated product description.
[0813] Step 4:
[0814] The data obtained from the generating AI is processed by the server device, converted into an appropriate format, and then transmitted to the terminal device. The input data is the output data from the generating AI (e.g., product description), and the output data is the data converted into a format for transmission to the terminal device.
[0815] Step 5:
[0816] The terminal device displays data received from the server to the user. The user confirms the generated data (e.g., product description, outfit suggestions, sales information). The input data is the generated data sent from the server, and the output data is the information displayed to the user.
[0817] Step 6:
[0818] The user reviews the displayed information and takes the next action. For example, they may request more detailed information or decide to purchase a product at a physical store. The input data is the user's action based on the generated information, and the output data is the user's next action entered into the terminal device.
[0819] Through the above processing steps, a system is realized in which data generated based on user requests is efficiently provided, thereby improving user convenience.
[0820] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0821] This invention combines an emotion engine with a multifunctional smartphone system using generative AI, adjusting the system's responses and display content based on the user's emotional state. This system consists of a "terminal means," a "server means," "generative AI technology," and an "emotion engine."
[0822] First, the user operates the terminal device and requests a specific task (e.g., practicing English conversation, suggesting cooking recipes, video editing). For example, if the user wants to practice English conversation, they use a dedicated app on the terminal device to input the request "Practice English conversation." Next, the terminal device sends this request to the server device. The server device receives the user's request and activates the generative AI and emotion engine. The emotion engine recognizes the user's emotions from their facial expressions and tone of voice, and based on this information, the generative AI generates an appropriate conversation scenario. This scenario is then sent back to the terminal device, and the user begins practicing English conversation accordingly.
[0823] This document describes an example of a cooking recipe suggestion system. When a user requests a cooking recipe, they enter "a dish using tomatoes and chicken" into a dedicated app on their terminal device. The terminal device sends the request to a server device, which uses a generation AI and an emotion engine to generate an optimal recipe based on the user's preferences, available ingredients, and emotional state. The generated recipe is then sent to the terminal device and displayed to the user. The user can then cook according to the recipe. If the user is tired, the emotion engine can recognize this and suggest a simpler recipe.
[0824] Let's explain a specific example of video editing. When a user wants to edit a video, they operate a terminal device to select the video they want to edit and request editing. The terminal device uploads the video material to a server device, which uses a generation AI and emotion engine to suggest editing options. For example, the server device can analyze the user's emotional state and suggest simple editing options that minimize stress. The suggestions are sent to the terminal device, and the user reviews them. If necessary, the user provides correction instructions, and the final editing is completed by the server device. The completed video is sent back to the terminal device for the user to review.
[0825] The process would look like this:
[0826] First, the user operates a terminal device and requests a specific task (English conversation, recipe suggestions, video editing). The terminal device sends the request to the server device, which activates the generative AI and emotion engine. The emotion engine recognizes the user's emotional state and passes this information to the generative AI to generate the optimal scenario and suggestions. The generated data is sent to the terminal device and displayed to the user. The user then performs the next action based on this data. At each step, the emotion engine recognizes the user's emotional state in real time and adjusts the interface and response content accordingly.
[0827] To realize this system, the server must be a computer system with high computing power and have software installed to run the generative AI and emotion engine. Furthermore, the terminal must have an interface for user operation, communication capabilities for communicating with the server, and also be equipped with a camera and microphone for emotion recognition.
[0828] Thus, by combining a generative AI and an emotion engine, the system of the present invention can more precisely address the diverse needs of users, and as a result, improve the user experience.
[0829] The following describes the processing flow.
[0830] English conversation practice
[0831] Step 1:
[0832] A user wants to practice English conversation and clicks the "Practice English Conversation" button on the app on their device.
[0833] Step 2:
[0834] The terminal receives the user's request and sends the request to the server.
[0835] Step 3:
[0836] The server receives the request and activates the generative AI and emotion engine.
[0837] Step 4:
[0838] The emotion engine analyzes the user's facial expressions and voice tone transmitted from the device to recognize the user's emotional state.
[0839] Step 5:
[0840] The server uses AI to generate conversation scenarios based on user information (past English conversation history and level) and emotional state.
[0841] Step 6:
[0842] The server generates a conversation scenario and sends it to the terminal device.
[0843] Step 7:
[0844] The device displays the conversation scenario it received to the user and presents the user with an initial question.
[0845] Step 8:
[0846] The user enters their answer to the question in text or voice.
[0847] Step 9:
[0848] The terminal transmits the user's response, along with their facial expressions and tone of voice during the response, to the server.
[0849] Step 10:
[0850] The server analyzes the user's responses and emotional state, and then uses AI to generate appropriate questions and feedback.
[0851] Step 11:
[0852] The server sends any new questions or feedback it generates to the terminal device.
[0853] Step 12:
[0854] The device displays new questions and feedback to the user.
[0855] Step 13:
[0856] The user responds again, and the process continues.
[0857] Cooking recipe suggestions
[0858] Step 1:
[0859] A user requests a cooking recipe and uses a dedicated app on their device to request a "tomato and chicken recipe."
[0860] Step 2:
[0861] The terminal receives the request and sends the request to the server.
[0862] Step 3:
[0863] The server receives the request and activates the generative AI and emotion engine.
[0864] Step 4:
[0865] The emotion engine analyzes the user's facial expressions and voice tone transmitted from the device to recognize the user's emotional state.
[0866] Step 5:
[0867] The server uses AI to generate the optimal recipe based on the user's preferences, available ingredients, and emotional state.
[0868] Step 6:
[0869] The server sends the generated recipe to the terminal device.
[0870] Step 7:
[0871] The device displays the received recipe to the user. For example, it displays a "Tomato and Chicken Cream Stew Recipe".
[0872] Step 8:
[0873] The user checks the displayed recipe and begins cooking.
[0874] Step 9:
[0875] The device displays instructions to the user for each step of the recipe.
[0876] Step 10:
[0877] The user follows instructions and proceeds with cooking. Their facial expressions and tone of voice during cooking are also analyzed by an emotion engine.
[0878] Video editing
[0879] Step 1:
[0880] The user wants to edit a video and uses their device to select the video they want to edit.
[0881] Step 2:
[0882] The device uploads the selected video material to the server.
[0883] Step 3:
[0884] The server receives the video footage and activates the generation AI and emotion engine.
[0885] Step 4:
[0886] The emotion engine analyzes the user's facial expressions and voice transmitted from the device to recognize the user's emotional state.
[0887] Step 5:
[0888] The server uses AI to generate editing suggestions (e.g., scene transition effects or music additions) based on the video content and the user's emotional state.
[0889] Step 6:
[0890] The server sends the generated editing suggestions to the terminal device.
[0891] Step 7:
[0892] The device displays editing suggestions to the user. For example, it might show "recommended scene transition effects" or "add music."
[0893] Step 8:
[0894] The user reviews the suggestion and instructs on any necessary changes. For example, they might instruct to "change the music."
[0895] Step 9:
[0896] The server receives user instructions, and the final editing work is performed by a generation AI.
[0897] Step 10:
[0898] The server renders the video and generates the final product.
[0899] Step 11:
[0900] The server sends the finished video to the terminal device.
[0901] Step 12:
[0902] The device displays the finished video to the user, who then reviews, saves, or shares it.
[0903] (Example 2)
[0904] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0905] Traditional systems only provided standard responses and suggestions to user requests, making it difficult to offer personalized services that took into account the user's emotional state. This resulted in problems such as decreased user satisfaction and a lower quality of experience.
[0906] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0907] In this invention, the server includes information terminal device means for receiving requests from users, computer means for providing data generated based on the requests, information terminal device means for displaying the generated data and providing an interface for the user to perform the next operation, emotion analysis means for recognizing the user's emotional state, and generation AI means for generating conversation scenarios and suggestions based on the emotional state. This makes it possible to provide highly personalized services that respond to the user's emotional state.
[0908] An "information terminal device means" is a device that a user operates and inputs requests into, and that is equipped with communication functions and a user interface for sending requests.
[0909] A "computer means" is a high-performance computer system that receives requests, processes data using generative AI technology and sentiment analysis technology, and provides the generated data.
[0910] "Emotional analysis means" refers to technologies and devices that analyze a user's facial expressions, tone of voice, and other emotional indicators to recognize the user's emotional state.
[0911] "Generative AI means" refers to technologies and devices that use generative AI models to generate scenarios and suggestions based on user requests.
[0912] An "interface" refers to a system that includes user interfaces and input devices for reviewing generated data and performing subsequent actions.
[0913] The present invention is a system that provides services and responses based on the user's emotional state. This system is composed of an information terminal device, a computer, an emotion analysis device, a generation AI device, and an interface.
[0914] When a user requests a specific task using the information terminal device, the information terminal device transmits this request to the computer device. The computer device processes the received request and activates the generation AI device and the emotion analysis device. The emotion analysis device analyzes emotional indicators such as the user's facial expressions and tone of voice to recognize the user's emotional state. Subsequently, the generation AI device generates optimal scenarios and suggestions based on the recognized emotional state. The generated scenarios and suggestions are then transmitted back to the information terminal device device by the computer device and displayed to the user.
[0915] The following specific hardware and software can be used in this invention:
[0916] Information terminal device means: For example, a smartphone (iPhone or Android device) has a dedicated app installed and a built-in camera and microphone.
[0917] Computing equipment: High-performance computer systems (e.g., AWS, Google Cloud) are installed with software for running generative AI models and sentiment analysis techniques.
[0918] Emotion analysis methods include facial recognition technology (e.g., Microsoft Azure's facial recognition API) and speech analysis technology (e.g., Google Cloud Speech-to-Text API).
[0919] Generative AI method: Use a generative AI model (e.g., GPT series) to generate scenarios and suggestions based on user requests.
[0920] Specific example
[0921] English conversation practice
[0922] 1. The user launches the dedicated app on their smartphone and enters "I want to practice English conversation."
[0923] 2. The information terminal device transmits this request to the computer.
[0924] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[0925] 4. The emotion analysis system analyzes the user's facial expressions and voice to recognize that they are relaxed.
[0926] 5. The generation AI generates a relaxed conversation scenario and sends it to the computer.
[0927] 6. The computer means transmits the generated scenario to the information terminal device means.
[0928] 7. The information terminal device displays the conversation scenario to the user.
[0929] 8. The user practices English conversation following the provided scenario.
[0930] Cooking recipe suggestions
[0931] 1. The user enters "I would like suggestions for dishes using tomatoes and chicken."
[0932] 2. The information terminal device transmits a request to the computer.
[0933] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[0934] 4. The emotion analysis tool analyzes the user's emotional state and, for example, recognizes that the user is slightly tired.
[0935] 5. The AI generates simple cooking recipes that require little physical effort and sends them to the computer.
[0936] 6. The computer means transmits the generated recipe to the information terminal device means.
[0937] 7. The information terminal device displays the recipe to the user.
[0938] 8. The user begins cooking according to the suggested recipe.
[0939] Video editing
[0940] 1. The user selects the video they want to edit and requests, "Easily edit this video."
[0941] 2. The information terminal device uploads the video to the computer and sends an editing request.
[0942] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[0943] 4. The emotion analysis tool analyzes the user's emotional state and recognizes that they are experiencing stress.
[0944] 5. The generation AI means generates simple editing suggestions and sends them to the computer means.
[0945] 6. The computer means transmits the generated editing proposal to the information terminal device means.
[0946] 7. The information terminal device displays editing suggestions to the user.
[0947] 8. Review the editing suggestions provided by the user and make any necessary corrections.
[0948] Examples of prompt statements include:
[0949] English conversation practice: "I want to practice speaking English."
[0950] Recipe suggestion: "Please suggest a dish using tomatoes and chicken."
[0951] Video editing request: "Please edit this video easily."
[0952] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0953] Step 1:
[0954] The user operates an information terminal device and requests a specific task.
[0955] Input: The user launches the dedicated app and enters a prompt message (e.g., "I want to practice English conversation").
[0956] Output: The information terminal device retrieves the user's request data.
[0957] Specific action: The user submits a request using an input form on their smartphone screen.
[0958] Step 2:
[0959] The terminal sends a request to the server.
[0960] Input: Retrieved request data.
[0961] Output: Sends the request data to the server.
[0962] Specific operation: The information terminal device sends request data to the server via an HTTP request. For example, it POSTs the request data to the endpoint.
[0963] Step 3:
[0964] The server processes the request and activates the generative AI and emotion analysis tools.
[0965] Input: Request data sent from the terminal.
[0966] Output: Requirements for emotion analysis and generation AI means.
[0967] Specific operation: The server analyzes the request data and selects the appropriate processing pipeline. It initializes and starts the generative AI and sentiment analysis tools.
[0968] Step 4:
[0969] The emotion analysis tool recognizes the user's emotional state.
[0970] Input: User's facial expression images and audio data sent to the server.
[0971] Output: Data indicating emotional state (e.g., relaxed, tense, tired).
[0972] Specific operation: The emotion analysis system analyzes the user's facial expressions and voice, and evaluates their emotional state using facial recognition technology and voice tone analysis technology.
[0973] Step 5:
[0974] The generation AI generates appropriate scenarios and suggestions.
[0975] Input: Request data, emotion state data.
[0976] Output: Generated scenarios and suggested data.
[0977] Specific operation: The generative AI uses models such as the GPT series to create optimal scenarios and suggestions by combining user requests and emotional states.
[0978] Step 6:
[0979] The server sends the generated scenarios and suggestions to the terminal.
[0980] Input: Scenarios and suggested data generated by the AI.
[0981] Output: Scenario and suggestion data sent to the terminal.
[0982] Specific operation: The server sends the generated scenarios and suggestions to the terminal as an HTTP response.
[0983] Step 7:
[0984] The device displays scenarios and suggestions it has received to the user.
[0985] Input: Scenario and suggestion data sent from the server.
[0986] Output: User screen displaying scenarios and suggestions.
[0987] Specific operation: The device analyzes data and displays it on the user interface. For example, it displays a conversation scenario in a chat format.
[0988] Step 8:
[0989] Users perform tasks according to scenarios and suggestions.
[0990] Input: Scenario and suggested data displayed on the terminal.
[0991] Output: User task execution status.
[0992] Specific actions: The user follows the presented scenarios and suggestions to perform the specified tasks (e.g., practicing English conversation, cooking).
[0993] Step 9:
[0994] At each step, the emotion analysis system recognizes the user's emotional state in real time and adjusts the interface and responses accordingly.
[0995] Input: User's current emotional state data.
[0996] Output: Updates to the user interface and response content.
[0997] Specific operation: The emotion analysis system monitors the emotional state in real time and adjusts the displayed content and responses as needed. For example, if it detects that the user is tired, it displays a message prompting a simple action.
[0998] By clearly specifying the concrete inputs, outputs, and actions at each step, the system's processing logic can be explained in detail.
[0999] (Application Example 2)
[1000] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1001] Traditional content delivery services lacked a mechanism to provide optimal content based on the user's emotional state, making it difficult to appropriately recommend content that users were looking for at that moment. This could lead to a lack of improvement in the user experience and a decrease in user satisfaction.
[1002] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1003] In this invention, the server includes terminal means for receiving requests from users, server means for receiving requests from the terminal means and providing data generated based on the requests, terminal means for displaying the generated data received from the server means and providing an interface for the user to perform the next operation, an emotion engine for analyzing the user's facial image and voice data and recognizing their emotional state, and generative AI technology for generating optimal content recommendations based on the output of the emotion engine. This makes it possible to recommend optimal content according to the user's emotional state.
[1004] A "terminal device" is a device used to receive requests from users.
[1005] A "server device" is a device that provides data generated based on a request received from a terminal device.
[1006] An "interface" is a mechanism that displays the generated data received from a server and allows the user to perform the following operations.
[1007] An "emotion engine" is a system that analyzes a user's facial image and voice data to recognize their emotional state.
[1008] "Generative AI technology" is an artificial intelligence technology that generates optimal content recommendations based on the output of an emotion engine.
[1009] "Content recommendation" refers to the suggestion of optimal content provided while taking the user's emotional state into consideration.
[1010] This invention is a system that combines generative AI technology and an emotion engine to recommend optimal content based on the user's emotional state. The system consists of multiple terminal means, server means, an emotion engine, and generative AI technology.
[1011] System program
[1012] The system basically operates in the following way:
[1013] 1. A device (such as a smartphone or tablet) receives requests from the user. This includes actions by the user to request content (such as playing a video or entering a search query).
[1014] 2. The terminal device acquires the user's facial image and voice, and collects data for emotion recognition. A digital camera and microphone are used in this process.
[1015] 3. The server receives requests and emotion recognition data from the terminal. It then analyzes the user's emotional state using an emotion engine. The emotion engine uses facial recognition software such as "DeepFace".
[1016] 4. The server uses generative AI technology based on the emotional state and request content to recommend the most suitable content. This process utilizes generative AI models such as the language model "GPT-3". For example, if a user is in a "happy mood" and is looking for recommended videos, this generative AI model will generate appropriate prompts and suggest content.
[1017] Hardware and software
[1018] Hardware:
[1019] The device must be equipped with a camera and a microphone. This allows for real-time analysis of the user's facial expressions and voice.
[1020] software:
[1021] The emotion engine uses facial recognition software such as "DeepFace." The generative AI technology uses large-scale language models such as "GPT-3."
[1022] Examples of specific cases and prompt statements
[1023] Let's consider what makes a user happy and what kind of video they would enjoy watching.
[1024] Specific example:
[1025] The terminal device captures the user's facial image, and the emotion engine recognizes the emotion "happy." This information is sent to the server, and the generating AI model recommends the most suitable content based on the following prompt message.
[1026] Example of a prompt:
[1027] "Please recommend videos that best capture the 'happiness' that users are experiencing."
[1028] This recommendation system enables personalized content delivery based on the user's emotional state, resulting in a more comfortable user experience.
[1029] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1030] Step 1:
[1031] The user operates a device such as a smartphone to enter a request (e.g., video recommendations tailored to their emotions). The request data is entered into the device and sent to the next step.
[1032] Step 2:
[1033] The terminal device captures the user's facial image and voice data using a camera and microphone. The captured facial image and voice data become input data for analysis by the emotion engine. Digital image processing and voice analysis are performed here.
[1034] Step 3:
[1035] The terminal device sends request data and facial image / audio data to the server device. The server device receives this data and performs emotion analysis using an emotion engine. The input data is the request and emotion analysis data, and the output data is the user's emotional state (e.g., "happy").
[1036] Step 4:
[1037] The server uses an emotion engine to analyze facial images and audio data to identify the user's emotional state. Specifically, it uses facial recognition software such as "DeepFace" to determine the emotional state. The data processing performed in this process involves emotion estimation through image and audio analysis.
[1038] Step 5:
[1039] Based on the emotion engine's output data (the user's emotional state), the server uses generative AI technology to generate optimal content recommendations. Specifically, it uses the language model "GPT-3" to generate prompt sentences that correspond to the user's emotions and then makes content recommendations based on those prompts. This involves both prompt sentence generation and the generation of recommendation data using a large-scale language model.
[1040] Step 6:
[1041] The server sends content recommendation data generated by generative AI technology to the terminal device. The terminal device receives the recommendation data and displays it to the user through its interface. The data processing performed here is UI rendering, which takes the recommendation data from the server and displays it in an easy-to-understand manner for the user.
[1042] Step 7:
[1043] The user reviews the displayed content recommendations and selects the next action (e.g., playing a video). This provides feedback that improves the user experience and leads to further requests. In this step, a loop of data collection and feedback based on user interaction is crucial.
[1044] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1045] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1046] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1047] [Third Embodiment]
[1048] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1049] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1050] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1051] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1052] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1053] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1054] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1055] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1056] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1057] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1058] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1059] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1060] This invention relates to a multifunctional smartphone system using generative AI, which enables users to efficiently perform various daily tasks. This system consists of "terminal means," "server means," and "generative AI technology."
[1061] First, the user operates the terminal device and requests a specific task (e.g., practicing English conversation, suggesting cooking recipes, video editing). For example, if the user wants to practice English conversation, they would use a dedicated app on the terminal device to input the request "Start practicing English conversation." Next, the terminal device sends this request to the server device. The server device receives the user's request and uses a generation AI to generate an appropriate conversation scenario. The generated conversation scenario is then sent back to the terminal device, and the user begins practicing English conversation accordingly.
[1062] Next, we will describe an example of a cooking recipe suggestion. When a user requests a cooking recipe, they input "a dish using tomatoes and chicken" using a dedicated app on their terminal device. The terminal device sends the request to the server device, which uses a generation AI to generate an optimal recipe based on the user's preferences and available ingredients. The generated recipe is then sent to the terminal device and displayed to the user. The user can then cook according to that recipe.
[1063] Next, we will explain a specific example of video editing. When a user wants to edit a video, they operate a terminal device to select the video they want to edit and request editing. The terminal device uploads the video material to a server device, and the server device uses AI generation to suggest editing options. For example, the server device might suggest "scene transition effects" or "adding music." The suggestions are sent to the terminal device, and the user reviews them. The user can then provide correction instructions as needed, and the final editing is completed by the server device. The completed video is sent back to the terminal device for the user to review.
[1064] To realize this system, the server is a computer system with high computing power and has software installed to execute generative AI technology. The terminal is equipped with an interface for user operation and has communication functions for communicating with the server.
[1065] The following is a specific example of a processing flow:
[1066] 1. The user operates a device and requests a specific task. (Examples: practicing English conversation, suggesting cooking recipes, video editing)
[1067] 2. The terminal sends the user's request to the server.
[1068] 3. The server receives the request and uses the generation AI to generate optimal data (e.g., conversation scenarios, cooking recipes, editing suggestions).
[1069] 4. The server transmits the generated data to the terminal device.
[1070] 5. The device displays the received data to the user and prompts them to take the next action.
[1071] 6. The user performs the following actions (e.g., responding to an English conversation, executing a cooking recipe, editing and correcting an edit).
[1072] Through the above process, a multi-functional smartphone system using generative AI can efficiently support the user's daily life.
[1073] The following describes the processing flow.
[1074] English conversation practice
[1075] Step 1:
[1076] The user wants to start practicing English conversation and operates the app on their device, clicking the "Practice English Conversation" button.
[1077] Step 2:
[1078] The terminal receives the user's request and sends the request to the server.
[1079] Step 3:
[1080] The server receives the request and activates the generation AI.
[1081] Step 4:
[1082] The server references user information (past English conversation history and level) and generates an appropriate conversation scenario.
[1083] Step 5:
[1084] The server generates a conversation scenario and sends it to the terminal device.
[1085] Step 6:
[1086] The device displays the conversation scenario it received to the user and presents the user with an initial question.
[1087] Step 7:
[1088] The user enters their answer to the question in text or voice.
[1089] Step 8:
[1090] The terminal sends the user's response to the server.
[1091] Step 9:
[1092] The server analyzes the user's responses and uses AI to generate appropriate next questions and feedback.
[1093] Step 10:
[1094] The server sends any new questions or feedback it generates to the terminal device.
[1095] Step 11:
[1096] The device displays new questions and feedback to the user.
[1097] Step 12:
[1098] The user responds again, and the process continues.
[1099] Cooking recipe suggestions
[1100] Step 1:
[1101] A user requests a cooking recipe and uses a dedicated app on their device to request a "tomato and chicken recipe."
[1102] Step 2:
[1103] The terminal receives the request and sends the request to the server.
[1104] Step 3:
[1105] The server receives the request and activates the generation AI.
[1106] Step 4:
[1107] The server collects information about the user's preferences and available ingredients to generate the optimal recipe.
[1108] Step 5:
[1109] The server sends the generated recipe to the terminal device.
[1110] Step 6:
[1111] The device displays the received recipe to the user.
[1112] Step 7:
[1113] The user views the displayed recipe and begins cooking.
[1114] Step 8:
[1115] The device displays instructions to the user for each step of the recipe.
[1116] Step 9:
[1117] The user follows the instructions and continues cooking.
[1118] Video editing
[1119] Step 1:
[1120] The user wants to edit a video and uses their device to select the video they want to edit.
[1121] Step 2:
[1122] The device uploads the selected video material to the server.
[1123] Step 3:
[1124] The server receives the video footage and activates the generation AI.
[1125] Step 4:
[1126] The server generates editing suggestions (e.g., scene transition effects or music additions) based on the video content.
[1127] Step 5:
[1128] The server sends the generated editing suggestions to the terminal device.
[1129] Step 6:
[1130] The device displays editing suggestions to the user.
[1131] Step 7:
[1132] The user reviews the proposal and indicates any necessary changes.
[1133] Step 8:
[1134] The server receives user instructions and performs the final editing.
[1135] Step 9:
[1136] The server renders the video and generates the final product.
[1137] Step 10:
[1138] The server sends the finished video to the terminal device.
[1139] Step 11:
[1140] The device displays the finished video to the user, who then reviews and saves it.
[1141] (Example 1)
[1142] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1143] Traditional smartphone applications lacked the automation of content generation based on user requests, resulting in significant user effort and time. Furthermore, they provided insufficient support for users to efficiently and effectively perform specific activities. In particular, traditional systems lacked flexibility and responsiveness for a wide range of tasks, such as practicing English conversation, getting cooking recipe suggestions, and video editing.
[1144] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1145] In this invention, the server includes a mobile terminal means for receiving requests from a user, a computer server means for receiving requests from the mobile terminal means and providing data generated based on the requests, a mobile terminal means for displaying the generated data received from the computer server means and providing an interface for the user to perform the next operation, a computer server means for generating optimal data based on the requests using a generation AI model, and a mobile terminal means for displaying specific instructions for the user to perform the next operation based on the generated data. This enables the user to efficiently perform a wide range of tasks, such as practicing English conversation, getting cooking recipe suggestions, and editing videos.
[1146] A "user" refers to a person who operates a mobile device to request a specific task.
[1147] "Mobile terminal means" refers to a portable electronic device that is operated by a user and has the function of communicating with a server.
[1148] A "computer server" refers to a high-performance computer system that has the function of processing user requests and providing generated data.
[1149] A "generative AI model" refers to artificial intelligence technology that generates optimal data based on user requests.
[1150] A "request" refers to a request for a specific task made by a user via a mobile device.
[1151] "Generated data" refers to information created by the computer server using a generation AI model, based on user requests.
[1152] "Interface" refers to the means of operation, such as screens and buttons, that users use to operate a mobile device.
[1153] Modes for carrying out the invention
[1154] This invention relates to a multifunctional smartphone system that enables users to efficiently perform a wide range of daily tasks. Specifically, it is a system that provides data generated based on requests from users, and is implemented using a mobile terminal, a computer server, and a generative AI model.
[1155] The system consists of the following main elements:
[1156] 1. Mobile device means:
[1157] The mobile device provides an interface for users to operate and enter requests. Users can enter various requests using a dedicated application. For example, requests such as "Start practicing English conversation," "Recipe a dish using tomatoes and chicken," or "Please give me suggestions for video editing" can be entered.
[1158] 2. Computer server means:
[1159] The computing server receives requests from users and generates optimal data using a generative AI model. The server has high computing power and utilizes generative AI models such as OpenAI GPT-3.5. The generated data varies depending on the content of the request. For example, conversation scenarios are generated for English conversation practice, optimal recipes for cooking recipe suggestions, and editing suggestions for video editing.
[1160] 3. Generative AI Models:
[1161] The generative AI model is responsible for generating content based on user requests. It takes text-based prompts as input and generates relevant information. For example, it processes prompts such as "Suggest the best recipe using tomatoes and chicken" or "Generate a scenario suitable for practicing English conversation."
[1162] Specific examples are given below.
[1163] English conversation practice
[1164] The user launches a dedicated app on their smartphone and enters "Start practicing English conversation." The device sends this request to a server, which uses a generative AI model to generate an appropriate conversation scenario. The generated scenario is sent back to the device, and the user begins practicing English conversation according to that scenario.
[1165] Example prompt: "Start practicing English conversation."
[1166] Cooking recipe suggestions
[1167] The user enters "a dish using tomatoes and chicken" into a dedicated app. The device sends the request to the server, which uses a generative AI model to generate the optimal recipe based on the user's preferences and available ingredients. The generated recipe is sent back to the device, and the user cooks according to that recipe.
[1168] Example prompt: "A dish using tomatoes and chicken"
[1169] Video editing suggestions
[1170] The user selects the video they want to edit using a dedicated app and requests "video editing suggestions." The device uploads the video footage to a server, which uses a generation AI model to generate editing suggestions (e.g., scene transition effects, adding music). The generated editing suggestions are sent back to the device, where the user reviews them and provides correction instructions as needed.
[1171] Example of a prompt: "Please edit this video and add scene transition effects."
[1172] As described above, the coordinated operation of each element enables the optimal generation and provision of data in response to user requests. This invention is a system that efficiently and effectively supports the user's daily life.
[1173] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1174] Steps for practicing English conversation
[1175] Step 1:
[1176] The user enters a request to "start practicing English conversation."
[1177] Input: User request ("Start practicing English conversation")
[1178] Specific operation: The user inputs data using a dedicated app on their mobile device.
[1179] Step 2:
[1180] The terminal receives user input and sends a request to the server.
[1181] Input: User Request
[1182] Output: Request data from terminal to server
[1183] Specific operation: The terminal converts the request into the appropriate format and sends it to the server over the internet.
[1184] Step 3:
[1185] The server receives the request and prompts the generated AI model.
[1186] Input: Request received by the server
[1187] Output: Prompt message for the generating AI model ("Generate a scenario suitable for practicing English conversation")
[1188] Specific operation: The server analyzes the request content and creates and inputs a prompt sentence suitable for the generated AI model.
[1189] Step 4:
[1190] The generative AI model generates appropriate conversation scenarios.
[1191] Input: Prompt message
[1192] Output: Generated conversation scenario
[1193] Specific operation: The generative AI model (e.g., OpenAI GPT-3.5) generates conversation scenarios based on prompts.
[1194] Step 5:
[1195] The server sends the generated conversation scenario to the terminal.
[1196] Input: Generated conversation scenario
[1197] Output: Conversation scenario data from server to terminal
[1198] Specific operation: The server sends the generated scenario to the terminal in the appropriate format.
[1199] Step 6:
[1200] The device displays a conversation scenario to the user.
[1201] Input: Conversation scenario data from the server
[1202] Output: Conversation scenario displayed on the user's device
[1203] Specific operation: The terminal renders the received scenario on the display screen.
[1204] Step 7:
[1205] The user practices English conversation based on the displayed conversation scenario.
[1206] Input: Displayed conversation scenario
[1207] Output: User's English conversation response
[1208] Specific operation: The user responds through the terminal.
[1209] Processing steps for suggesting cooking recipes
[1210] Step 1:
[1211] The user enters a request for "a dish using tomatoes and chicken."
[1212] Input: User request ("A dish using tomatoes and chicken")
[1213] Specific operation: The user inputs data using a dedicated app on their mobile device.
[1214] Step 2:
[1215] The terminal receives user input and sends a request to the server.
[1216] Input: User Request
[1217] Output: Request data from terminal to server
[1218] Specific operation: The terminal converts the request into the appropriate format and sends it to the server over the internet.
[1219] Step 3:
[1220] The server receives the request and prompts the generated AI model.
[1221] Input: Request received by the server
[1222] Output: Prompt message for the generative AI model ("Suggest the best recipe using tomatoes and chicken as ingredients")
[1223] Specific operation: The server analyzes the request content and creates and inputs a prompt sentence suitable for the generated AI model.
[1224] Step 4:
[1225] The generative AI model generates the optimal cooking recipe.
[1226] Input: Prompt message
[1227] Output: Generated cooking recipe
[1228] Specific operation: The generative AI model generates cooking recipes based on prompts.
[1229] Step 5:
[1230] The server sends the generated cooking recipe to the terminal.
[1231] Input: Generated cooking recipe
[1232] Output: Recipe data from server to terminal
[1233] Specific operation: The server sends the generated recipe to the terminal in the appropriate format.
[1234] Step 6:
[1235] The device displays cooking recipes to the user.
[1236] Input: Recipe data from the server
[1237] Output: Recipe displayed on the user's device
[1238] Specific action: The device renders the received recipe on the display screen.
[1239] Step 7:
[1240] The user cooks according to the displayed recipe.
[1241] Input: Displayed recipe
[1242] Output: Cooked food
[1243] Specific actions: The user follows the displayed recipe and cooks the dish.
[1244] Processing steps for video editing suggestions
[1245] Step 1:
[1246] The user selects the video they want to edit and requests "video editing suggestions."
[1247] Input: User request ("I would like suggestions for video editing")
[1248] Specific operation: The user inputs data using a dedicated app on their mobile device.
[1249] Step 2:
[1250] The device uploads video footage to the server.
[1251] Input: User selected video footage
[1252] Output: Video data from terminal to server
[1253] Specific operation: The device compresses the video footage and sends it to the server via the internet.
[1254] Step 3:
[1255] The server receives the video footage and inputs prompts into the generating AI model.
[1256] Input: Video footage received by the server
[1257] Output: Prompt message for the generating AI model ("Generate editing suggestions suitable for the video")
[1258] Specific operation: The server analyzes the video material and creates and inputs prompt sentences suitable for the generating AI model.
[1259] Step 4:
[1260] The generative AI model generates appropriate editing suggestions.
[1261] Input: Prompt message
[1262] Output: Generated editing suggestions
[1263] Specific operation: The generation AI model generates editing suggestions (e.g., scene transition effects, adding music) based on prompts.
[1264] Step 5:
[1265] The server sends the generated editing suggestions to the terminal.
[1266] Input: Generated editing suggestions
[1267] Output: Editing suggestion data from server to terminal
[1268] Specific operation: The server sends the generated editing suggestions to the terminal in the appropriate format.
[1269] Step 6:
[1270] The device displays editing suggestions to the user.
[1271] Input: Editing suggestion data from the server
[1272] Output: Editing suggestions displayed on the user's device
[1273] Specific action: The terminal renders the received editing suggestions on the display screen.
[1274] Step 7:
[1275] The user reviews the displayed editing suggestions and provides correction instructions as needed.
[1276] Input: Displayed editing suggestions
[1277] Output: Correction instructions
[1278] Specific actions: The user reviews the displayed editing suggestions and enters correction instructions into the terminal as needed.
[1279] Step 8:
[1280] The server performs the final editing and sends the completed video to the device.
[1281] Input: Correction Instructions
[1282] Output: Completed video
[1283] Specific operation: The server performs the final editing based on the correction instructions and sends the completed video to the terminal.
[1284] Step 9:
[1285] The device displays the completed video to the user.
[1286] Input: Completed video data from the server
[1287] Output: The completed video will be displayed on the user's device.
[1288] Specific operation: The device renders the received completed video on the display screen and allows the user to confirm it.
[1289] (Application Example 1)
[1290] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1291] Modern consumers and store staff are required to efficiently handle a wide range of tasks, including providing real-time product information, suggesting optimal outfit combinations, and instantly announcing sales information. However, fulfilling these requirements with a single system is difficult and time-consuming, resulting in significant costs. This invention aims to improve convenience for both users and store staff by efficiently providing diverse information services required in physical stores using generative AI technology.
[1292] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1293] In this invention, the server includes terminal device means for receiving requests from users, server device means for receiving requests from the terminal device means and providing data generated based on the requests, and terminal device means for displaying the generated data received from the server device means and providing an interface for the user to perform subsequent operations. This makes it possible to provide product descriptions, coordination suggestions, and sales information.
[1294] A "user" is a person who uses a system to obtain information or request tasks.
[1295] A "terminal device means" is an electronic device used to receive requests from users or to display data from a server.
[1296] "Server device means" refers to a computer system that processes and provides data generated based on requests received from terminal device means.
[1297] "Generative AI technology" is artificial intelligence technology that generates optimal data and suggestions based on user requests.
[1298] "Product description" refers to providing users and store staff with detailed information about a specific product.
[1299] "Coordination suggestions" refer to providing optimal fashion and styling ideas by combining specific items.
[1300] "Sale information" refers to data about special prices and discounts currently available.
[1301] An "interface" is an operation screen or input / output means that allows a user to interact with a system and perform the following operations.
[1302] "Feedback" is the process of providing responses to user requests and related information.
[1303] The system for carrying out this invention comprises a terminal device means for receiving requests from users, a server device means for generating and providing data based on the requests, and a terminal device means for displaying the generated data to the user and providing an interface for performing subsequent operations. In this system, AI generation technology is utilized to generate optimal data in response to requests.
[1304] Explanation of the program's processing
[1305] 1. Terminal device means:
[1306] The terminal device is an electronic device such as a smartphone or tablet that receives requests from the user. The user uses a dedicated application on the terminal device to input specific tasks (e.g., product description, outfit suggestions, sales information inquiries). For example, the user inputs a request such as, "Please tell me more about the XYZ smartwatch."
[1307] 2. Server device means:
[1308] The server device receives requests sent from the frontend and generates appropriate data using generative AI technology. The server system includes a computer system with high computing power and has software installed to run generative AI models (e.g., OpenAI API). When a user request is sent to the server, the server passes a prompt message to the generative AI and sends the obtained result back to the terminal device.
[1309] For example, if a user requests, "Please suggest an outfit using a red dress, black boots, and a white coat," the server will pass this prompt to the AI that generates the appropriate outfit suggestions.
[1310] 3. Means of providing the interface:
[1311] The terminal device provides an interface that displays the generated data received from the server to the user. The user can then review the generated data and perform the following actions: for example, review the generated product descriptions and styling suggestions, make more detailed requests as needed, or consider purchasing items at a physical store based on the provided information.
[1312] Example prompt statements
[1313] Product Description:
[1314] "Could you please provide a detailed product description for the XYZ smartwatch?"
[1315] Outfit suggestions:
[1316] "Please suggest outfit ideas using a red dress, black boots, and a white coat."
[1317] This enables the provision of information that facilitates efficient communication between users and store staff in physical stores. Advanced information processing is achieved by utilizing generative AI models.
[1318] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1319] Step 1:
[1320] The user operates the terminal device and enters a specific request through a dedicated application. For example, if the user enters "Please tell me more about the XYZ smartwatch," this request is temporarily stored in the terminal device. The input data is saved as the request content and sent to the next processing step.
[1321] Step 2:
[1322] The terminal device transmits the received user request to the server device. The transmitted data is the user's request content (e.g., a request for product description), which is converted to an appropriate format and sent to the server. The input data is the user's request content, and the output data is the request data sent to the server.
[1323] Step 3:
[1324] The server device analyzes the received request and generates appropriate data using generative AI technology. The generative AI model constructs a prompt sentence based on the request content and sends it to the generative AI. The input data is the user's request content, which is processed into a prompt sentence. The output data obtained from the generative AI is specific information such as the generated product description.
[1325] Step 4:
[1326] The data obtained from the generating AI is processed by the server device, converted into an appropriate format, and then transmitted to the terminal device. The input data is the output data from the generating AI (e.g., product description), and the output data is the data converted into a format for transmission to the terminal device.
[1327] Step 5:
[1328] The terminal device displays data received from the server to the user. The user confirms the generated data (e.g., product description, outfit suggestions, sales information). The input data is the generated data sent from the server, and the output data is the information displayed to the user.
[1329] Step 6:
[1330] The user reviews the displayed information and takes the next action. For example, they may request more detailed information or decide to purchase a product at a physical store. The input data is the user's action based on the generated information, and the output data is the user's next action entered into the terminal device.
[1331] Through the above processing steps, a system is realized in which data generated based on user requests is efficiently provided, thereby improving user convenience.
[1332] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1333] This invention combines an emotion engine with a multifunctional smartphone system using generative AI, adjusting the system's responses and display content based on the user's emotional state. This system consists of a "terminal means," a "server means," "generative AI technology," and an "emotion engine."
[1334] First, the user operates the terminal device and requests a specific task (e.g., practicing English conversation, suggesting cooking recipes, video editing). For example, if the user wants to practice English conversation, they use a dedicated app on the terminal device to input the request "Practice English conversation." Next, the terminal device sends this request to the server device. The server device receives the user's request and activates the generative AI and emotion engine. The emotion engine recognizes the user's emotions from their facial expressions and tone of voice, and based on this information, the generative AI generates an appropriate conversation scenario. This scenario is then sent back to the terminal device, and the user begins practicing English conversation accordingly.
[1335] This document describes an example of a cooking recipe suggestion system. When a user requests a cooking recipe, they enter "a dish using tomatoes and chicken" into a dedicated app on their terminal device. The terminal device sends the request to a server device, which uses a generation AI and an emotion engine to generate an optimal recipe based on the user's preferences, available ingredients, and emotional state. The generated recipe is then sent to the terminal device and displayed to the user. The user can then cook according to the recipe. If the user is tired, the emotion engine can recognize this and suggest a simpler recipe.
[1336] Let's explain a specific example of video editing. When a user wants to edit a video, they operate a terminal device to select the video they want to edit and request editing. The terminal device uploads the video material to a server device, which uses a generation AI and emotion engine to suggest editing options. For example, the server device can analyze the user's emotional state and suggest simple editing options that minimize stress. The suggestions are sent to the terminal device, and the user reviews them. If necessary, the user provides correction instructions, and the final editing is completed by the server device. The completed video is sent back to the terminal device for the user to review.
[1337] The process would look like this:
[1338] First, the user operates a terminal device and requests a specific task (English conversation, recipe suggestions, video editing). The terminal device sends the request to the server device, which activates the generative AI and emotion engine. The emotion engine recognizes the user's emotional state and passes this information to the generative AI to generate the optimal scenario and suggestions. The generated data is sent to the terminal device and displayed to the user. The user then performs the next action based on this data. At each step, the emotion engine recognizes the user's emotional state in real time and adjusts the interface and response content accordingly.
[1339] To realize this system, the server must be a computer system with high computing power and have software installed to run the generative AI and emotion engine. Furthermore, the terminal must have an interface for user operation, communication capabilities for communicating with the server, and also be equipped with a camera and microphone for emotion recognition.
[1340] Thus, by combining a generative AI and an emotion engine, the system of the present invention can more precisely address the diverse needs of users, and as a result, improve the user experience.
[1341] The following describes the processing flow.
[1342] English conversation practice
[1343] Step 1:
[1344] A user wants to practice English conversation and clicks the "Practice English Conversation" button on the app on their device.
[1345] Step 2:
[1346] The terminal receives the user's request and sends the request to the server.
[1347] Step 3:
[1348] The server receives the request and activates the generative AI and emotion engine.
[1349] Step 4:
[1350] The emotion engine analyzes the user's facial expressions and voice tone transmitted from the device to recognize the user's emotional state.
[1351] Step 5:
[1352] The server uses AI to generate conversation scenarios based on user information (past English conversation history and level) and emotional state.
[1353] Step 6:
[1354] The server generates a conversation scenario and sends it to the terminal device.
[1355] Step 7:
[1356] The device displays the conversation scenario it received to the user and presents the user with an initial question.
[1357] Step 8:
[1358] The user enters their answer to the question in text or voice.
[1359] Step 9:
[1360] The terminal transmits the user's response, along with their facial expressions and tone of voice during the response, to the server.
[1361] Step 10:
[1362] The server analyzes the user's responses and emotional state, and then uses AI to generate appropriate questions and feedback.
[1363] Step 11:
[1364] The server sends any new questions or feedback it generates to the terminal device.
[1365] Step 12:
[1366] The device displays new questions and feedback to the user.
[1367] Step 13:
[1368] The user responds again, and the process continues.
[1369] Cooking recipe suggestions
[1370] Step 1:
[1371] A user requests a cooking recipe and uses a dedicated app on their device to request a "tomato and chicken recipe."
[1372] Step 2:
[1373] The terminal receives the request and sends the request to the server.
[1374] Step 3:
[1375] The server receives the request and activates the generative AI and emotion engine.
[1376] Step 4:
[1377] The emotion engine analyzes the user's facial expressions and voice tone transmitted from the device to recognize the user's emotional state.
[1378] Step 5:
[1379] The server uses AI to generate the optimal recipe based on the user's preferences, available ingredients, and emotional state.
[1380] Step 6:
[1381] The server sends the generated recipe to the terminal device.
[1382] Step 7:
[1383] The device displays the received recipe to the user. For example, it displays a "Tomato and Chicken Cream Stew Recipe".
[1384] Step 8:
[1385] The user checks the displayed recipe and begins cooking.
[1386] Step 9:
[1387] The device displays instructions to the user for each step of the recipe.
[1388] Step 10:
[1389] The user follows instructions and proceeds with cooking. Their facial expressions and tone of voice during cooking are also analyzed by an emotion engine.
[1390] Video editing
[1391] Step 1:
[1392] The user wants to edit a video and uses their device to select the video they want to edit.
[1393] Step 2:
[1394] The device uploads the selected video material to the server.
[1395] Step 3:
[1396] The server receives the video footage and activates the generation AI and emotion engine.
[1397] Step 4:
[1398] The emotion engine analyzes the user's facial expressions and voice transmitted from the device to recognize the user's emotional state.
[1399] Step 5:
[1400] The server uses AI to generate editing suggestions (e.g., scene transition effects or music additions) based on the video content and the user's emotional state.
[1401] Step 6:
[1402] The server sends the generated editing suggestions to the terminal device.
[1403] Step 7:
[1404] The device displays editing suggestions to the user. For example, it might show "recommended scene transition effects" or "add music."
[1405] Step 8:
[1406] The user reviews the suggestion and instructs on any necessary changes. For example, they might instruct to "change the music."
[1407] Step 9:
[1408] The server receives user instructions, and the final editing work is performed by a generation AI.
[1409] Step 10:
[1410] The server renders the video and generates the final product.
[1411] Step 11:
[1412] The server sends the finished video to the terminal device.
[1413] Step 12:
[1414] The device displays the finished video to the user, who then reviews, saves, or shares it.
[1415] (Example 2)
[1416] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1417] Traditional systems only provided standard responses and suggestions to user requests, making it difficult to offer personalized services that took into account the user's emotional state. This resulted in problems such as decreased user satisfaction and a lower quality of experience.
[1418] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1419] In this invention, the server includes information terminal device means for receiving requests from users, computer means for providing data generated based on the requests, information terminal device means for displaying the generated data and providing an interface for the user to perform the next operation, emotion analysis means for recognizing the user's emotional state, and generation AI means for generating conversation scenarios and suggestions based on the emotional state. This makes it possible to provide highly personalized services that respond to the user's emotional state.
[1420] An "information terminal device means" is a device that a user operates and inputs requests into, and that is equipped with communication functions and a user interface for sending requests.
[1421] A "computer means" is a high-performance computer system that receives requests, processes data using generative AI technology and sentiment analysis technology, and provides the generated data.
[1422] "Emotional analysis means" refers to technologies and devices that analyze a user's facial expressions, tone of voice, and other emotional indicators to recognize the user's emotional state.
[1423] "Generative AI means" refers to technologies and devices that use generative AI models to generate scenarios and suggestions based on user requests.
[1424] An "interface" refers to a system that includes user interfaces and input devices for reviewing generated data and performing subsequent actions.
[1425] The present invention is a system that provides services and responses based on the user's emotional state. This system is composed of an information terminal device, a computer, an emotion analysis device, a generation AI device, and an interface.
[1426] When a user requests a specific task using the information terminal device, the information terminal device transmits this request to the computer device. The computer device processes the received request and activates the generation AI device and the emotion analysis device. The emotion analysis device analyzes emotional indicators such as the user's facial expressions and tone of voice to recognize the user's emotional state. Subsequently, the generation AI device generates optimal scenarios and suggestions based on the recognized emotional state. The generated scenarios and suggestions are then transmitted back to the information terminal device device by the computer device and displayed to the user.
[1427] The following specific hardware and software can be used in this invention:
[1428] Information terminal device means: For example, a smartphone (iPhone or Android device) has a dedicated app installed and a built-in camera and microphone.
[1429] Computing equipment: High-performance computer systems (e.g., AWS, Google Cloud) are installed with software for running generative AI models and sentiment analysis techniques.
[1430] Emotion analysis methods include facial recognition technology (e.g., Microsoft Azure's facial recognition API) and speech analysis technology (e.g., Google Cloud Speech-to-Text API).
[1431] Generative AI method: Use a generative AI model (e.g., GPT series) to generate scenarios and suggestions based on user requests.
[1432] Specific example
[1433] English conversation practice
[1434] 1. The user launches the dedicated app on their smartphone and enters "I want to practice English conversation."
[1435] 2. The information terminal device transmits this request to the computer.
[1436] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[1437] 4. The emotion analysis system analyzes the user's facial expressions and voice to recognize that they are relaxed.
[1438] 5. The generation AI generates a relaxed conversation scenario and sends it to the computer.
[1439] 6. The computer means transmits the generated scenario to the information terminal device means.
[1440] 7. The information terminal device displays the conversation scenario to the user.
[1441] 8. The user practices English conversation following the provided scenario.
[1442] Cooking recipe suggestions
[1443] 1. The user enters "I would like suggestions for dishes using tomatoes and chicken."
[1444] 2. The information terminal device transmits a request to the computer.
[1445] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[1446] 4. The emotion analysis tool analyzes the user's emotional state and, for example, recognizes that the user is slightly tired.
[1447] 5. The AI generates simple cooking recipes that require little physical effort and sends them to the computer.
[1448] 6. The computer means transmits the generated recipe to the information terminal device means.
[1449] 7. The information terminal device displays the recipe to the user.
[1450] 8. The user begins cooking according to the suggested recipe.
[1451] Video editing
[1452] 1. The user selects the video they want to edit and requests, "Easily edit this video."
[1453] 2. The information terminal device uploads the video to the computer and sends an editing request.
[1454] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[1455] 4. The emotion analysis tool analyzes the user's emotional state and recognizes that they are experiencing stress.
[1456] 5. The generation AI means generates simple editing suggestions and sends them to the computer means.
[1457] 6. The computer means transmits the generated editing proposal to the information terminal device means.
[1458] 7. The information terminal device displays editing suggestions to the user.
[1459] 8. Review the editing suggestions provided by the user and make any necessary corrections.
[1460] Examples of prompt statements include:
[1461] English conversation practice: "I want to practice speaking English."
[1462] Recipe suggestion: "Please suggest a dish using tomatoes and chicken."
[1463] Video editing request: "Please edit this video easily."
[1464] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1465] Step 1:
[1466] The user operates an information terminal device and requests a specific task.
[1467] Input: The user launches the dedicated app and enters a prompt message (e.g., "I want to practice English conversation").
[1468] Output: The information terminal device retrieves the user's request data.
[1469] Specific action: The user submits a request using an input form on their smartphone screen.
[1470] Step 2:
[1471] The terminal sends a request to the server.
[1472] Input: Retrieved request data.
[1473] Output: Sends the request data to the server.
[1474] Specific operation: The information terminal device sends request data to the server via an HTTP request. For example, it POSTs the request data to the endpoint.
[1475] Step 3:
[1476] The server processes the request and activates the generative AI and emotion analysis tools.
[1477] Input: Request data sent from the terminal.
[1478] Output: Requirements for emotion analysis and generation AI means.
[1479] Specific operation: The server analyzes the request data and selects the appropriate processing pipeline. It initializes and starts the generative AI and sentiment analysis tools.
[1480] Step 4:
[1481] The emotion analysis tool recognizes the user's emotional state.
[1482] Input: User's facial expression images and audio data sent to the server.
[1483] Output: Data indicating emotional state (e.g., relaxed, tense, tired).
[1484] Specific operation: The emotion analysis system analyzes the user's facial expressions and voice, and evaluates their emotional state using facial recognition technology and voice tone analysis technology.
[1485] Step 5:
[1486] The generation AI generates appropriate scenarios and suggestions.
[1487] Input: Request data, emotion state data.
[1488] Output: Generated scenarios and suggested data.
[1489] Specific operation: The generative AI uses models such as the GPT series to create optimal scenarios and suggestions by combining user requests and emotional states.
[1490] Step 6:
[1491] The server sends the generated scenarios and suggestions to the terminal.
[1492] Input: Scenarios and suggested data generated by the AI.
[1493] Output: Scenario and suggestion data sent to the terminal.
[1494] Specific operation: The server sends the generated scenarios and suggestions to the terminal as an HTTP response.
[1495] Step 7:
[1496] The device displays scenarios and suggestions it has received to the user.
[1497] Input: Scenario and suggestion data sent from the server.
[1498] Output: User screen displaying scenarios and suggestions.
[1499] Specific operation: The device analyzes data and displays it on the user interface. For example, it displays a conversation scenario in a chat format.
[1500] Step 8:
[1501] Users perform tasks according to scenarios and suggestions.
[1502] Input: Scenario and suggested data displayed on the terminal.
[1503] Output: User task execution status.
[1504] Specific actions: The user follows the presented scenarios and suggestions to perform the specified tasks (e.g., practicing English conversation, cooking).
[1505] Step 9:
[1506] At each step, the emotion analysis system recognizes the user's emotional state in real time and adjusts the interface and responses accordingly.
[1507] Input: User's current emotional state data.
[1508] Output: Updates to the user interface and response content.
[1509] Specific operation: The emotion analysis system monitors the emotional state in real time and adjusts the displayed content and responses as needed. For example, if it detects that the user is tired, it displays a message prompting a simple action.
[1510] By clearly specifying the concrete inputs, outputs, and actions at each step, the system's processing logic can be explained in detail.
[1511] (Application Example 2)
[1512] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1513] Traditional content delivery services lacked a mechanism to provide optimal content based on the user's emotional state, making it difficult to appropriately recommend content that users were looking for at that moment. This could lead to a lack of improvement in the user experience and a decrease in user satisfaction.
[1514] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1515] In this invention, the server includes terminal means for receiving requests from users, server means for receiving requests from the terminal means and providing data generated based on the requests, terminal means for displaying the generated data received from the server means and providing an interface for the user to perform the next operation, an emotion engine for analyzing the user's facial image and voice data and recognizing their emotional state, and generative AI technology for generating optimal content recommendations based on the output of the emotion engine. This makes it possible to recommend optimal content according to the user's emotional state.
[1516] A "terminal device" is a device used to receive requests from users.
[1517] A "server device" is a device that provides data generated based on a request received from a terminal device.
[1518] An "interface" is a mechanism that displays the generated data received from a server and allows the user to perform the following operations.
[1519] An "emotion engine" is a system that analyzes a user's facial image and voice data to recognize their emotional state.
[1520] "Generative AI technology" is an artificial intelligence technology that generates optimal content recommendations based on the output of an emotion engine.
[1521] "Content recommendation" refers to the suggestion of optimal content provided while taking the user's emotional state into consideration.
[1522] This invention is a system that combines generative AI technology and an emotion engine to recommend optimal content based on the user's emotional state. The system consists of multiple terminal means, server means, an emotion engine, and generative AI technology.
[1523] System program
[1524] The system basically operates in the following way:
[1525] 1. A device (such as a smartphone or tablet) receives requests from the user. This includes actions by the user to request content (such as playing a video or entering a search query).
[1526] 2. The terminal device acquires the user's facial image and voice, and collects data for emotion recognition. A digital camera and microphone are used in this process.
[1527] 3. The server receives requests and emotion recognition data from the terminal. It then analyzes the user's emotional state using an emotion engine. The emotion engine uses facial recognition software such as "DeepFace".
[1528] 4. The server uses generative AI technology based on the emotional state and request content to recommend the most suitable content. This process utilizes generative AI models such as the language model "GPT-3". For example, if a user is in a "happy mood" and is looking for recommended videos, this generative AI model will generate appropriate prompts and suggest content.
[1529] Hardware and software
[1530] Hardware:
[1531] The device must be equipped with a camera and a microphone. This allows for real-time analysis of the user's facial expressions and voice.
[1532] software:
[1533] The emotion engine uses facial recognition software such as "DeepFace." The generative AI technology uses large-scale language models such as "GPT-3."
[1534] Examples of specific cases and prompt statements
[1535] Let's consider what makes a user happy and what kind of video they would enjoy watching.
[1536] Specific example:
[1537] The terminal device captures the user's facial image, and the emotion engine recognizes the emotion "happy." This information is sent to the server, and the generating AI model recommends the most suitable content based on the following prompt message.
[1538] Example of a prompt:
[1539] "Please recommend videos that best capture the 'happiness' that users are experiencing."
[1540] This recommendation system enables personalized content delivery based on the user's emotional state, resulting in a more comfortable user experience.
[1541] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1542] Step 1:
[1543] The user operates a device such as a smartphone to enter a request (e.g., video recommendations tailored to their emotions). The request data is entered into the device and sent to the next step.
[1544] Step 2:
[1545] The terminal device captures the user's facial image and voice data using a camera and microphone. The captured facial image and voice data become input data for analysis by the emotion engine. Digital image processing and voice analysis are performed here.
[1546] Step 3:
[1547] The terminal device sends request data and facial image / audio data to the server device. The server device receives this data and performs emotion analysis using an emotion engine. The input data is the request and emotion analysis data, and the output data is the user's emotional state (e.g., "happy").
[1548] Step 4:
[1549] The server uses an emotion engine to analyze facial images and audio data to identify the user's emotional state. Specifically, it uses facial recognition software such as "DeepFace" to determine the emotional state. The data processing performed in this process involves emotion estimation through image and audio analysis.
[1550] Step 5:
[1551] Based on the emotion engine's output data (the user's emotional state), the server uses generative AI technology to generate optimal content recommendations. Specifically, it uses the language model "GPT-3" to generate prompt sentences that correspond to the user's emotions and then makes content recommendations based on those prompts. This involves both prompt sentence generation and the generation of recommendation data using a large-scale language model.
[1552] Step 6:
[1553] The server sends content recommendation data generated by generative AI technology to the terminal device. The terminal device receives the recommendation data and displays it to the user through its interface. The data processing performed here is UI rendering, which takes the recommendation data from the server and displays it in an easy-to-understand manner for the user.
[1554] Step 7:
[1555] The user reviews the displayed content recommendations and selects the next action (e.g., playing a video). This provides feedback that improves the user experience and leads to further requests. In this step, a loop of data collection and feedback based on user interaction is crucial.
[1556] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1557] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1558] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1559] [Fourth Embodiment]
[1560] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1561] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1562] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1563] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1564] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1565] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1566] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1567] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1568] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1569] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1570] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1571] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1572] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1573] This invention relates to a multifunctional smartphone system using generative AI, which enables users to efficiently perform various daily tasks. This system consists of "terminal means," "server means," and "generative AI technology."
[1574] First, the user operates the terminal device and requests a specific task (e.g., practicing English conversation, suggesting cooking recipes, video editing). For example, if the user wants to practice English conversation, they would use a dedicated app on the terminal device to input the request "Start practicing English conversation." Next, the terminal device sends this request to the server device. The server device receives the user's request and uses a generation AI to generate an appropriate conversation scenario. The generated conversation scenario is then sent back to the terminal device, and the user begins practicing English conversation accordingly.
[1575] Next, we will describe an example of a cooking recipe suggestion. When a user requests a cooking recipe, they input "a dish using tomatoes and chicken" using a dedicated app on their terminal device. The terminal device sends the request to the server device, which uses a generation AI to generate an optimal recipe based on the user's preferences and available ingredients. The generated recipe is then sent to the terminal device and displayed to the user. The user can then cook according to that recipe.
[1576] Next, we will explain a specific example of video editing. When a user wants to edit a video, they operate a terminal device to select the video they want to edit and request editing. The terminal device uploads the video material to a server device, and the server device uses AI generation to suggest editing options. For example, the server device might suggest "scene transition effects" or "adding music." The suggestions are sent to the terminal device, and the user reviews them. The user can then provide correction instructions as needed, and the final editing is completed by the server device. The completed video is sent back to the terminal device for the user to review.
[1577] To realize this system, the server is a computer system with high computing power and has software installed to execute generative AI technology. The terminal is equipped with an interface for user operation and has communication functions for communicating with the server.
[1578] The following is a specific example of a processing flow:
[1579] 1. The user operates a device and requests a specific task. (Examples: practicing English conversation, suggesting cooking recipes, video editing)
[1580] 2. The terminal sends the user's request to the server.
[1581] 3. The server receives the request and uses the generation AI to generate optimal data (e.g., conversation scenarios, cooking recipes, editing suggestions).
[1582] 4. The server transmits the generated data to the terminal device.
[1583] 5. The device displays the received data to the user and prompts them to take the next action.
[1584] 6. The user performs the following actions (e.g., responding to an English conversation, executing a cooking recipe, editing and correcting an edit).
[1585] Through the above process, a multi-functional smartphone system using generative AI can efficiently support the user's daily life.
[1586] The following describes the processing flow.
[1587] English conversation practice
[1588] Step 1:
[1589] The user wants to start practicing English conversation and operates the app on their device, clicking the "Practice English Conversation" button.
[1590] Step 2:
[1591] The terminal receives the user's request and sends the request to the server.
[1592] Step 3:
[1593] The server receives the request and activates the generation AI.
[1594] Step 4:
[1595] The server references user information (past English conversation history and level) and generates an appropriate conversation scenario.
[1596] Step 5:
[1597] The server generates a conversation scenario and sends it to the terminal device.
[1598] Step 6:
[1599] The device displays the conversation scenario it received to the user and presents the user with an initial question.
[1600] Step 7:
[1601] The user enters their answer to the question in text or voice.
[1602] Step 8:
[1603] The terminal sends the user's response to the server.
[1604] Step 9:
[1605] The server analyzes the user's responses and uses AI to generate appropriate next questions and feedback.
[1606] Step 10:
[1607] The server sends any new questions or feedback it generates to the terminal device.
[1608] Step 11:
[1609] The device displays new questions and feedback to the user.
[1610] Step 12:
[1611] The user responds again, and the process continues.
[1612] Cooking recipe suggestions
[1613] Step 1:
[1614] A user requests a cooking recipe and uses a dedicated app on their device to request a "tomato and chicken recipe."
[1615] Step 2:
[1616] The terminal receives the request and sends the request to the server.
[1617] Step 3:
[1618] The server receives the request and activates the generation AI.
[1619] Step 4:
[1620] The server collects information about the user's preferences and available ingredients to generate the optimal recipe.
[1621] Step 5:
[1622] The server sends the generated recipe to the terminal device.
[1623] Step 6:
[1624] The device displays the received recipe to the user.
[1625] Step 7:
[1626] The user views the displayed recipe and begins cooking.
[1627] Step 8:
[1628] The device displays instructions to the user for each step of the recipe.
[1629] Step 9:
[1630] The user follows the instructions and continues cooking.
[1631] Video editing
[1632] Step 1:
[1633] The user wants to edit a video and uses their device to select the video they want to edit.
[1634] Step 2:
[1635] The device uploads the selected video material to the server.
[1636] Step 3:
[1637] The server receives the video footage and activates the generation AI.
[1638] Step 4:
[1639] The server generates editing suggestions (e.g., scene transition effects or music additions) based on the video content.
[1640] Step 5:
[1641] The server sends the generated editing suggestions to the terminal device.
[1642] Step 6:
[1643] The device displays editing suggestions to the user.
[1644] Step 7:
[1645] The user reviews the proposal and indicates any necessary changes.
[1646] Step 8:
[1647] The server receives user instructions and performs the final editing.
[1648] Step 9:
[1649] The server renders the video and generates the final product.
[1650] Step 10:
[1651] The server sends the finished video to the terminal device.
[1652] Step 11:
[1653] The device displays the finished video to the user, who then reviews and saves it.
[1654] (Example 1)
[1655] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1656] Traditional smartphone applications lacked the automation of content generation based on user requests, resulting in significant user effort and time. Furthermore, they provided insufficient support for users to efficiently and effectively perform specific activities. In particular, traditional systems lacked flexibility and responsiveness for a wide range of tasks, such as practicing English conversation, getting cooking recipe suggestions, and video editing.
[1657] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1658] In this invention, the server includes a mobile terminal means for receiving requests from a user, a computer server means for receiving requests from the mobile terminal means and providing data generated based on the requests, a mobile terminal means for displaying the generated data received from the computer server means and providing an interface for the user to perform the next operation, a computer server means for generating optimal data based on the requests using a generation AI model, and a mobile terminal means for displaying specific instructions for the user to perform the next operation based on the generated data. This enables the user to efficiently perform a wide range of tasks, such as practicing English conversation, getting cooking recipe suggestions, and editing videos.
[1659] A "user" refers to a person who operates a mobile device to request a specific task.
[1660] "Mobile terminal means" refers to a portable electronic device that is operated by a user and has the function of communicating with a server.
[1661] A "computer server" refers to a high-performance computer system that has the function of processing user requests and providing generated data.
[1662] A "generative AI model" refers to artificial intelligence technology that generates optimal data based on user requests.
[1663] A "request" refers to a request for a specific task made by a user via a mobile device.
[1664] "Generated data" refers to information created by the computer server using a generation AI model, based on user requests.
[1665] "Interface" refers to the means of operation, such as screens and buttons, that users use to operate a mobile device.
[1666] Modes for carrying out the invention
[1667] This invention relates to a multifunctional smartphone system that enables users to efficiently perform a wide range of daily tasks. Specifically, it is a system that provides data generated based on requests from users, and is implemented using a mobile terminal, a computer server, and a generative AI model.
[1668] The system consists of the following main elements:
[1669] 1. Mobile device means:
[1670] The mobile device provides an interface for users to operate and enter requests. Users can enter various requests using a dedicated application. For example, requests such as "Start practicing English conversation," "Recipe a dish using tomatoes and chicken," or "Please give me suggestions for video editing" can be entered.
[1671] 2. Computer server means:
[1672] The computing server receives requests from users and generates optimal data using a generative AI model. The server has high computing power and utilizes generative AI models such as OpenAI GPT-3.5. The generated data varies depending on the content of the request. For example, conversation scenarios are generated for English conversation practice, optimal recipes for cooking recipe suggestions, and editing suggestions for video editing.
[1673] 3. Generative AI Models:
[1674] The generative AI model is responsible for generating content based on user requests. It takes text-based prompts as input and generates relevant information. For example, it processes prompts such as "Suggest the best recipe using tomatoes and chicken" or "Generate a scenario suitable for practicing English conversation."
[1675] Specific examples are given below.
[1676] English conversation practice
[1677] The user launches a dedicated app on their smartphone and enters "Start practicing English conversation." The device sends this request to a server, which uses a generative AI model to generate an appropriate conversation scenario. The generated scenario is sent back to the device, and the user begins practicing English conversation according to that scenario.
[1678] Example prompt: "Start practicing English conversation."
[1679] Cooking recipe suggestions
[1680] The user enters "a dish using tomatoes and chicken" into a dedicated app. The device sends the request to the server, which uses a generative AI model to generate the optimal recipe based on the user's preferences and available ingredients. The generated recipe is sent back to the device, and the user cooks according to that recipe.
[1681] Example prompt: "A dish using tomatoes and chicken"
[1682] Video editing suggestions
[1683] The user selects the video they want to edit using a dedicated app and requests "video editing suggestions." The device uploads the video footage to a server, which uses a generation AI model to generate editing suggestions (e.g., scene transition effects, adding music). The generated editing suggestions are sent back to the device, where the user reviews them and provides correction instructions as needed.
[1684] Example of a prompt: "Please edit this video and add scene transition effects."
[1685] As described above, the coordinated operation of each element enables the optimal generation and provision of data in response to user requests. This invention is a system that efficiently and effectively supports the user's daily life.
[1686] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1687] Steps for practicing English conversation
[1688] Step 1:
[1689] The user enters a request to "start practicing English conversation."
[1690] Input: User request ("Start practicing English conversation")
[1691] Specific operation: The user inputs data using a dedicated app on their mobile device.
[1692] Step 2:
[1693] The terminal receives user input and sends a request to the server.
[1694] Input: User Request
[1695] Output: Request data from terminal to server
[1696] Specific operation: The terminal converts the request into the appropriate format and sends it to the server over the internet.
[1697] Step 3:
[1698] The server receives the request and prompts the generated AI model.
[1699] Input: Request received by the server
[1700] Output: Prompt message for the generating AI model ("Generate a scenario suitable for practicing English conversation")
[1701] Specific operation: The server analyzes the request content and creates and inputs a prompt sentence suitable for the generated AI model.
[1702] Step 4:
[1703] The generative AI model generates appropriate conversation scenarios.
[1704] Input: Prompt message
[1705] Output: Generated conversation scenario
[1706] Specific operation: The generative AI model (e.g., OpenAI GPT-3.5) generates conversation scenarios based on prompts.
[1707] Step 5:
[1708] The server sends the generated conversation scenario to the terminal.
[1709] Input: Generated conversation scenario
[1710] Output: Conversation scenario data from server to terminal
[1711] Specific operation: The server sends the generated scenario to the terminal in the appropriate format.
[1712] Step 6:
[1713] The device displays a conversation scenario to the user.
[1714] Input: Conversation scenario data from the server
[1715] Output: Conversation scenario displayed on the user's device
[1716] Specific operation: The terminal renders the received scenario on the display screen.
[1717] Step 7:
[1718] The user practices English conversation based on the displayed conversation scenario.
[1719] Input: Displayed conversation scenario
[1720] Output: User's English conversation response
[1721] Specific operation: The user responds through the terminal.
[1722] Processing steps for suggesting cooking recipes
[1723] Step 1:
[1724] The user enters a request for "a dish using tomatoes and chicken."
[1725] Input: User request ("A dish using tomatoes and chicken")
[1726] Specific operation: The user inputs data using a dedicated app on their mobile device.
[1727] Step 2:
[1728] The terminal receives user input and sends a request to the server.
[1729] Input: User Request
[1730] Output: Request data from terminal to server
[1731] Specific operation: The terminal converts the request into the appropriate format and sends it to the server over the internet.
[1732] Step 3:
[1733] The server receives the request and prompts the generated AI model.
[1734] Input: Request received by the server
[1735] Output: Prompt message for the generative AI model ("Suggest the best recipe using tomatoes and chicken as ingredients")
[1736] Specific operation: The server analyzes the request content and creates and inputs a prompt sentence suitable for the generated AI model.
[1737] Step 4:
[1738] The generative AI model generates the optimal cooking recipe.
[1739] Input: Prompt message
[1740] Output: Generated cooking recipe
[1741] Specific operation: The generative AI model generates cooking recipes based on prompts.
[1742] Step 5:
[1743] The server sends the generated cooking recipe to the terminal.
[1744] Input: Generated cooking recipe
[1745] Output: Recipe data from server to terminal
[1746] Specific operation: The server sends the generated recipe to the terminal in the appropriate format.
[1747] Step 6:
[1748] The device displays cooking recipes to the user.
[1749] Input: Recipe data from the server
[1750] Output: Recipe displayed on the user's device
[1751] Specific action: The device renders the received recipe on the display screen.
[1752] Step 7:
[1753] The user cooks according to the displayed recipe.
[1754] Input: Displayed recipe
[1755] Output: Cooked food
[1756] Specific actions: The user follows the displayed recipe and cooks the dish.
[1757] Processing steps for video editing suggestions
[1758] Step 1:
[1759] The user selects the video they want to edit and requests "video editing suggestions."
[1760] Input: User request ("I would like suggestions for video editing")
[1761] Specific operation: The user inputs data using a dedicated app on their mobile device.
[1762] Step 2:
[1763] The device uploads video footage to the server.
[1764] Input: User selected video footage
[1765] Output: Video data from terminal to server
[1766] Specific operation: The device compresses the video footage and sends it to the server via the internet.
[1767] Step 3:
[1768] The server receives the video footage and inputs prompts into the generating AI model.
[1769] Input: Video footage received by the server
[1770] Output: Prompt message for the generating AI model ("Generate editing suggestions suitable for the video")
[1771] Specific operation: The server analyzes the video material and creates and inputs prompt sentences suitable for the generating AI model.
[1772] Step 4:
[1773] The generative AI model generates appropriate editing suggestions.
[1774] Input: Prompt message
[1775] Output: Generated editing suggestions
[1776] Specific operation: The generation AI model generates editing suggestions (e.g., scene transition effects, adding music) based on prompts.
[1777] Step 5:
[1778] The server sends the generated editing suggestions to the terminal.
[1779] Input: Generated editing suggestions
[1780] Output: Editing suggestion data from server to terminal
[1781] Specific operation: The server sends the generated editing suggestions to the terminal in the appropriate format.
[1782] Step 6:
[1783] The device displays editing suggestions to the user.
[1784] Input: Editing suggestion data from the server
[1785] Output: Editing suggestions displayed on the user's device
[1786] Specific action: The terminal renders the received editing suggestions on the display screen.
[1787] Step 7:
[1788] The user reviews the displayed editing suggestions and provides correction instructions as needed.
[1789] Input: Displayed editing suggestions
[1790] Output: Correction instructions
[1791] Specific actions: The user reviews the displayed editing suggestions and enters correction instructions into the terminal as needed.
[1792] Step 8:
[1793] The server performs the final editing and sends the completed video to the device.
[1794] Input: Correction Instructions
[1795] Output: Completed video
[1796] Specific operation: The server performs the final editing based on the correction instructions and sends the completed video to the terminal.
[1797] Step 9:
[1798] The device displays the completed video to the user.
[1799] Input: Completed video data from the server
[1800] Output: The completed video will be displayed on the user's device.
[1801] Specific operation: The device renders the received completed video on the display screen and allows the user to confirm it.
[1802] (Application Example 1)
[1803] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1804] Modern consumers and store staff are required to efficiently handle a wide range of tasks, including providing real-time product information, suggesting optimal outfit combinations, and instantly announcing sales information. However, fulfilling these requirements with a single system is difficult and time-consuming, resulting in significant costs. This invention aims to improve convenience for both users and store staff by efficiently providing diverse information services required in physical stores using generative AI technology.
[1805] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1806] In this invention, the server includes terminal device means for receiving requests from users, server device means for receiving requests from the terminal device means and providing data generated based on the requests, and terminal device means for displaying the generated data received from the server device means and providing an interface for the user to perform subsequent operations. This makes it possible to provide product descriptions, coordination suggestions, and sales information.
[1807] A "user" is a person who uses a system to obtain information or request tasks.
[1808] A "terminal device means" is an electronic device used to receive requests from users or to display data from a server.
[1809] "Server device means" refers to a computer system that processes and provides data generated based on requests received from terminal device means.
[1810] "Generative AI technology" is artificial intelligence technology that generates optimal data and suggestions based on user requests.
[1811] "Product description" refers to providing users and store staff with detailed information about a specific product.
[1812] "Coordination suggestions" refer to providing optimal fashion and styling ideas by combining specific items.
[1813] "Sale information" refers to data about special prices and discounts currently available.
[1814] An "interface" is an operation screen or input / output means that allows a user to interact with a system and perform the following operations.
[1815] "Feedback" is the process of providing responses to user requests and related information.
[1816] The system for carrying out this invention comprises a terminal device means for receiving requests from users, a server device means for generating and providing data based on the requests, and a terminal device means for displaying the generated data to the user and providing an interface for performing subsequent operations. In this system, AI generation technology is utilized to generate optimal data in response to requests.
[1817] Explanation of the program's processing
[1818] 1. Terminal device means:
[1819] The terminal device is an electronic device such as a smartphone or tablet that receives requests from the user. The user uses a dedicated application on the terminal device to input specific tasks (e.g., product description, outfit suggestions, sales information inquiries). For example, the user inputs a request such as, "Please tell me more about the XYZ smartwatch."
[1820] 2. Server device means:
[1821] The server device receives requests sent from the frontend and generates appropriate data using generative AI technology. The server system includes a computer system with high computing power and has software installed to run generative AI models (e.g., OpenAI API). When a user request is sent to the server, the server passes a prompt message to the generative AI and sends the obtained result back to the terminal device.
[1822] For example, if a user requests, "Please suggest an outfit using a red dress, black boots, and a white coat," the server will pass this prompt to the AI that generates the appropriate outfit suggestions.
[1823] 3. Means of providing the interface:
[1824] The terminal device provides an interface that displays the generated data received from the server to the user. The user can then review the generated data and perform the following actions: for example, review the generated product descriptions and styling suggestions, make more detailed requests as needed, or consider purchasing items at a physical store based on the provided information.
[1825] Example prompt statements
[1826] Product Description:
[1827] "Could you please provide a detailed product description for the XYZ smartwatch?"
[1828] Outfit suggestions:
[1829] "Please suggest outfit ideas using a red dress, black boots, and a white coat."
[1830] This enables the provision of information that facilitates efficient communication between users and store staff in physical stores. Advanced information processing is achieved by utilizing generative AI models.
[1831] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1832] Step 1:
[1833] The user operates the terminal device and enters a specific request through a dedicated application. For example, if the user enters "Please tell me more about the XYZ smartwatch," this request is temporarily stored in the terminal device. The input data is saved as the request content and sent to the next processing step.
[1834] Step 2:
[1835] The terminal device transmits the received user request to the server device. The transmitted data is the user's request content (e.g., a request for product description), which is converted to an appropriate format and sent to the server. The input data is the user's request content, and the output data is the request data sent to the server.
[1836] Step 3:
[1837] The server device analyzes the received request and generates appropriate data using generative AI technology. The generative AI model constructs a prompt sentence based on the request content and sends it to the generative AI. The input data is the user's request content, which is processed into a prompt sentence. The output data obtained from the generative AI is specific information such as the generated product description.
[1838] Step 4:
[1839] The data obtained from the generating AI is processed by the server device, converted into an appropriate format, and then transmitted to the terminal device. The input data is the output data from the generating AI (e.g., product description), and the output data is the data converted into a format for transmission to the terminal device.
[1840] Step 5:
[1841] The terminal device displays data received from the server to the user. The user confirms the generated data (e.g., product description, outfit suggestions, sales information). The input data is the generated data sent from the server, and the output data is the information displayed to the user.
[1842] Step 6:
[1843] The user reviews the displayed information and takes the next action. For example, they may request more detailed information or decide to purchase a product at a physical store. The input data is the user's action based on the generated information, and the output data is the user's next action entered into the terminal device.
[1844] Through the above processing steps, a system is realized in which data generated based on user requests is efficiently provided, thereby improving user convenience.
[1845] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1846] This invention combines an emotion engine with a multifunctional smartphone system using generative AI, adjusting the system's responses and display content based on the user's emotional state. This system consists of a "terminal means," a "server means," "generative AI technology," and an "emotion engine."
[1847] First, the user operates the terminal device and requests a specific task (e.g., practicing English conversation, suggesting cooking recipes, video editing). For example, if the user wants to practice English conversation, they use a dedicated app on the terminal device to input the request "Practice English conversation." Next, the terminal device sends this request to the server device. The server device receives the user's request and activates the generative AI and emotion engine. The emotion engine recognizes the user's emotions from their facial expressions and tone of voice, and based on this information, the generative AI generates an appropriate conversation scenario. This scenario is then sent back to the terminal device, and the user begins practicing English conversation accordingly.
[1848] This document describes an example of a cooking recipe suggestion system. When a user requests a cooking recipe, they enter "a dish using tomatoes and chicken" into a dedicated app on their terminal device. The terminal device sends the request to a server device, which uses a generation AI and an emotion engine to generate an optimal recipe based on the user's preferences, available ingredients, and emotional state. The generated recipe is then sent to the terminal device and displayed to the user. The user can then cook according to the recipe. If the user is tired, the emotion engine can recognize this and suggest a simpler recipe.
[1849] Let's explain a specific example of video editing. When a user wants to edit a video, they operate a terminal device to select the video they want to edit and request editing. The terminal device uploads the video material to a server device, which uses a generation AI and emotion engine to suggest editing options. For example, the server device can analyze the user's emotional state and suggest simple editing options that minimize stress. The suggestions are sent to the terminal device, and the user reviews them. If necessary, the user provides correction instructions, and the final editing is completed by the server device. The completed video is sent back to the terminal device for the user to review.
[1850] The process would look like this:
[1851] First, the user operates a terminal device and requests a specific task (English conversation, recipe suggestions, video editing). The terminal device sends the request to the server device, which activates the generative AI and emotion engine. The emotion engine recognizes the user's emotional state and passes this information to the generative AI to generate the optimal scenario and suggestions. The generated data is sent to the terminal device and displayed to the user. The user then performs the next action based on this data. At each step, the emotion engine recognizes the user's emotional state in real time and adjusts the interface and response content accordingly.
[1852] To realize this system, the server must be a computer system with high computing power and have software installed to run the generative AI and emotion engine. Furthermore, the terminal must have an interface for user operation, communication capabilities for communicating with the server, and also be equipped with a camera and microphone for emotion recognition.
[1853] Thus, by combining a generative AI and an emotion engine, the system of the present invention can more precisely address the diverse needs of users, and as a result, improve the user experience.
[1854] The following describes the processing flow.
[1855] English conversation practice
[1856] Step 1:
[1857] A user wants to practice English conversation and clicks the "Practice English Conversation" button on the app on their device.
[1858] Step 2:
[1859] The terminal receives the user's request and sends the request to the server.
[1860] Step 3:
[1861] The server receives the request and activates the generative AI and emotion engine.
[1862] Step 4:
[1863] The emotion engine analyzes the user's facial expressions and voice tone transmitted from the device to recognize the user's emotional state.
[1864] Step 5:
[1865] The server uses AI to generate conversation scenarios based on user information (past English conversation history and level) and emotional state.
[1866] Step 6:
[1867] The server generates a conversation scenario and sends it to the terminal device.
[1868] Step 7:
[1869] The device displays the conversation scenario it received to the user and presents the user with an initial question.
[1870] Step 8:
[1871] The user enters their answer to the question in text or voice.
[1872] Step 9:
[1873] The terminal transmits the user's response, along with their facial expressions and tone of voice during the response, to the server.
[1874] Step 10:
[1875] The server analyzes the user's responses and emotional state, and then uses AI to generate appropriate questions and feedback.
[1876] Step 11:
[1877] The server sends any new questions or feedback it generates to the terminal device.
[1878] Step 12:
[1879] The device displays new questions and feedback to the user.
[1880] Step 13:
[1881] The user responds again, and the process continues.
[1882] Cooking recipe suggestions
[1883] Step 1:
[1884] A user requests a cooking recipe and uses a dedicated app on their device to request a "tomato and chicken recipe."
[1885] Step 2:
[1886] The terminal receives the request and sends the request to the server.
[1887] Step 3:
[1888] The server receives the request and activates the generative AI and emotion engine.
[1889] Step 4:
[1890] The emotion engine analyzes the user's facial expressions and voice tone transmitted from the device to recognize the user's emotional state.
[1891] Step 5:
[1892] The server uses AI to generate the optimal recipe based on the user's preferences, available ingredients, and emotional state.
[1893] Step 6:
[1894] The server sends the generated recipe to the terminal device.
[1895] Step 7:
[1896] The device displays the received recipe to the user. For example, it displays a "Tomato and Chicken Cream Stew Recipe".
[1897] Step 8:
[1898] The user checks the displayed recipe and begins cooking.
[1899] Step 9:
[1900] The device displays instructions to the user for each step of the recipe.
[1901] Step 10:
[1902] The user follows instructions and proceeds with cooking. Their facial expressions and tone of voice during cooking are also analyzed by an emotion engine.
[1903] Video editing
[1904] Step 1:
[1905] The user wants to edit a video and uses their device to select the video they want to edit.
[1906] Step 2:
[1907] The device uploads the selected video material to the server.
[1908] Step 3:
[1909] The server receives the video footage and activates the generation AI and emotion engine.
[1910] Step 4:
[1911] The emotion engine analyzes the user's facial expressions and voice transmitted from the device to recognize the user's emotional state.
[1912] Step 5:
[1913] The server uses AI to generate editing suggestions (e.g., scene transition effects or music additions) based on the video content and the user's emotional state.
[1914] Step 6:
[1915] The server sends the generated editing suggestions to the terminal device.
[1916] Step 7:
[1917] The device displays editing suggestions to the user. For example, it might show "recommended scene transition effects" or "add music."
[1918] Step 8:
[1919] The user reviews the suggestion and instructs on any necessary changes. For example, they might instruct to "change the music."
[1920] Step 9:
[1921] The server receives user instructions, and the final editing work is performed by a generation AI.
[1922] Step 10:
[1923] The server renders the video and generates the final product.
[1924] Step 11:
[1925] The server sends the finished video to the terminal device.
[1926] Step 12:
[1927] The device displays the finished video to the user, who then reviews, saves, or shares it.
[1928] (Example 2)
[1929] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1930] Traditional systems only provided standard responses and suggestions to user requests, making it difficult to offer personalized services that took into account the user's emotional state. This resulted in problems such as decreased user satisfaction and a lower quality of experience.
[1931] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1932] In this invention, the server includes information terminal device means for receiving requests from users, computer means for providing data generated based on the requests, information terminal device means for displaying the generated data and providing an interface for the user to perform the next operation, emotion analysis means for recognizing the user's emotional state, and generation AI means for generating conversation scenarios and suggestions based on the emotional state. This makes it possible to provide highly personalized services that respond to the user's emotional state.
[1933] An "information terminal device means" is a device that a user operates and inputs requests into, and that is equipped with communication functions and a user interface for sending requests.
[1934] A "computer means" is a high-performance computer system that receives requests, processes data using generative AI technology and sentiment analysis technology, and provides the generated data.
[1935] "Emotional analysis means" refers to technologies and devices that analyze a user's facial expressions, tone of voice, and other emotional indicators to recognize the user's emotional state.
[1936] "Generative AI means" refers to technologies and devices that use generative AI models to generate scenarios and suggestions based on user requests.
[1937] An "interface" refers to a system that includes user interfaces and input devices for reviewing generated data and performing subsequent actions.
[1938] The present invention is a system that provides services and responses based on the user's emotional state. This system is composed of an information terminal device, a computer, an emotion analysis device, a generation AI device, and an interface.
[1939] When a user requests a specific task using the information terminal device, the information terminal device transmits this request to the computer device. The computer device processes the received request and activates the generation AI device and the emotion analysis device. The emotion analysis device analyzes emotional indicators such as the user's facial expressions and tone of voice to recognize the user's emotional state. Subsequently, the generation AI device generates optimal scenarios and suggestions based on the recognized emotional state. The generated scenarios and suggestions are then transmitted back to the information terminal device device by the computer device and displayed to the user.
[1940] The following specific hardware and software can be used in this invention:
[1941] Information terminal device means: For example, a smartphone (iPhone or Android device) has a dedicated app installed and a built-in camera and microphone.
[1942] Computing equipment: High-performance computer systems (e.g., AWS, Google Cloud) are installed with software for running generative AI models and sentiment analysis techniques.
[1943] Emotion analysis methods include facial recognition technology (e.g., Microsoft Azure's facial recognition API) and speech analysis technology (e.g., Google Cloud Speech-to-Text API).
[1944] Generative AI method: Use a generative AI model (e.g., GPT series) to generate scenarios and suggestions based on user requests.
[1945] Specific example
[1946] English conversation practice
[1947] 1. The user launches the dedicated app on their smartphone and enters "I want to practice English conversation."
[1948] 2. The information terminal device transmits this request to the computer.
[1949] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[1950] 4. The emotion analysis system analyzes the user's facial expressions and voice to recognize that they are relaxed.
[1951] 5. The generation AI generates a relaxed conversation scenario and sends it to the computer.
[1952] 6. The computer means transmits the generated scenario to the information terminal device means.
[1953] 7. The information terminal device displays the conversation scenario to the user.
[1954] 8. The user practices English conversation following the provided scenario.
[1955] Cooking recipe suggestions
[1956] 1. The user enters "I would like suggestions for dishes using tomatoes and chicken."
[1957] 2. The information terminal device transmits a request to the computer.
[1958] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[1959] 4. The emotion analysis tool analyzes the user's emotional state and, for example, recognizes that the user is slightly tired.
[1960] 5. The AI generates simple cooking recipes that require little physical effort and sends them to the computer.
[1961] 6. The computer means transmits the generated recipe to the information terminal device means.
[1962] 7. The information terminal device displays the recipe to the user.
[1963] 8. The user begins cooking according to the suggested recipe.
[1964] Video editing
[1965] 1. The user selects the video they want to edit and requests, "Easily edit this video."
[1966] 2. The information terminal device uploads the video to the computer and sends an editing request.
[1967] 3. The computing means receives the request and activates the generation AI means and the emotion analysis means.
[1968] 4. The emotion analysis tool analyzes the user's emotional state and recognizes that they are experiencing stress.
[1969] 5. The generation AI means generates simple editing suggestions and sends them to the computer means.
[1970] 6. The computer means transmits the generated editing proposal to the information terminal device means.
[1971] 7. The information terminal device displays editing suggestions to the user.
[1972] 8. Review the editing suggestions provided by the user and make any necessary corrections.
[1973] Examples of prompt statements include:
[1974] English conversation practice: "I want to practice speaking English."
[1975] Recipe suggestion: "Please suggest a dish using tomatoes and chicken."
[1976] Video editing request: "Please edit this video easily."
[1977] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1978] Step 1:
[1979] The user operates an information terminal device and requests a specific task.
[1980] Input: The user launches the dedicated app and enters a prompt message (e.g., "I want to practice English conversation").
[1981] Output: The information terminal device retrieves the user's request data.
[1982] Specific action: The user submits a request using an input form on their smartphone screen.
[1983] Step 2:
[1984] The terminal sends a request to the server.
[1985] Input: Retrieved request data.
[1986] Output: Sends the request data to the server.
[1987] Specific operation: The information terminal device sends request data to the server via an HTTP request. For example, it POSTs the request data to the endpoint.
[1988] Step 3:
[1989] The server processes the request and activates the generative AI and emotion analysis tools.
[1990] Input: Request data sent from the terminal.
[1991] Output: Requirements for emotion analysis and generation AI means.
[1992] Specific operation: The server analyzes the request data and selects the appropriate processing pipeline. It initializes and starts the generative AI and sentiment analysis tools.
[1993] Step 4:
[1994] The emotion analysis tool recognizes the user's emotional state.
[1995] Input: User's facial expression images and audio data sent to the server.
[1996] Output: Data indicating emotional state (e.g., relaxed, tense, tired).
[1997] Specific operation: The emotion analysis system analyzes the user's facial expressions and voice, and evaluates their emotional state using facial recognition technology and voice tone analysis technology.
[1998] Step 5:
[1999] The generation AI generates appropriate scenarios and suggestions.
[2000] Input: Request data, emotion state data.
[2001] Output: Generated scenarios and suggested data.
[2002] Specific operation: The generative AI uses models such as the GPT series to create optimal scenarios and suggestions by combining user requests and emotional states.
[2003] Step 6:
[2004] The server sends the generated scenarios and suggestions to the terminal.
[2005] Input: Scenarios and suggested data generated by the AI.
[2006] Output: Scenario and suggestion data sent to the terminal.
[2007] Specific operation: The server sends the generated scenarios and suggestions to the terminal as an HTTP response.
[2008] Step 7:
[2009] The device displays scenarios and suggestions it has received to the user.
[2010] Input: Scenario and suggestion data sent from the server.
[2011] Output: User screen displaying scenarios and suggestions.
[2012] Specific operation: The device analyzes data and displays it on the user interface. For example, it displays a conversation scenario in a chat format.
[2013] Step 8:
[2014] Users perform tasks according to scenarios and suggestions.
[2015] Input: Scenario and suggested data displayed on the terminal.
[2016] Output: User task execution status.
[2017] Specific actions: The user follows the presented scenarios and suggestions to perform the specified tasks (e.g., practicing English conversation, cooking).
[2018] Step 9:
[2019] At each step, the emotion analysis system recognizes the user's emotional state in real time and adjusts the interface and responses accordingly.
[2020] Input: User's current emotional state data.
[2021] Output: Updates to the user interface and response content.
[2022] Specific operation: The emotion analysis system monitors the emotional state in real time and adjusts the displayed content and responses as needed. For example, if it detects that the user is tired, it displays a message prompting a simple action.
[2023] By clearly specifying the concrete inputs, outputs, and actions at each step, the system's processing logic can be explained in detail.
[2024] (Application Example 2)
[2025] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2026] Traditional content delivery services lacked a mechanism to provide optimal content based on the user's emotional state, making it difficult to appropriately recommend content that users were looking for at that moment. This could lead to a lack of improvement in the user experience and a decrease in user satisfaction.
[2027] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[2028] In this invention, the server includes terminal means for receiving requests from users, server means for receiving requests from the terminal means and providing data generated based on the requests, terminal means for displaying the generated data received from the server means and providing an interface for the user to perform the next operation, an emotion engine for analyzing the user's facial image and voice data and recognizing their emotional state, and generative AI technology for generating optimal content recommendations based on the output of the emotion engine. This makes it possible to recommend optimal content according to the user's emotional state.
[2029] A "terminal device" is a device used to receive requests from users.
[2030] A "server device" is a device that provides data generated based on a request received from a terminal device.
[2031] An "interface" is a mechanism that displays the generated data received from a server and allows the user to perform the following operations.
[2032] An "emotion engine" is a system that analyzes a user's facial image and voice data to recognize their emotional state.
[2033] "Generative AI technology" is an artificial intelligence technology that generates optimal content recommendations based on the output of an emotion engine.
[2034] "Content recommendation" refers to the suggestion of optimal content provided while taking the user's emotional state into consideration.
[2035] This invention is a system that combines generative AI technology and an emotion engine to recommend optimal content based on the user's emotional state. The system consists of multiple terminal means, server means, an emotion engine, and generative AI technology.
[2036] System program
[2037] The system basically operates in the following way:
[2038] 1. A device (such as a smartphone or tablet) receives requests from the user. This includes actions by the user to request content (such as playing a video or entering a search query).
[2039] 2. The terminal device acquires the user's facial image and voice, and collects data for emotion recognition. A digital camera and microphone are used in this process.
[2040] 3. The server receives requests and emotion recognition data from the terminal. It then analyzes the user's emotional state using an emotion engine. The emotion engine uses facial recognition software such as "DeepFace".
[2041] 4. The server uses generative AI technology based on the emotional state and request content to recommend the most suitable content. This process utilizes generative AI models such as the language model "GPT-3". For example, if a user is in a "happy mood" and is looking for recommended videos, this generative AI model will generate appropriate prompts and suggest content.
[2042] Hardware and software
[2043] Hardware:
[2044] The device must be equipped with a camera and a microphone. This allows for real-time analysis of the user's facial expressions and voice.
[2045] software:
[2046] The emotion engine uses facial recognition software such as "DeepFace." The generative AI technology uses large-scale language models such as "GPT-3."
[2047] Examples of specific cases and prompt statements
[2048] Let's consider what makes a user happy and what kind of video they would enjoy watching.
[2049] Specific example:
[2050] The terminal device captures the user's facial image, and the emotion engine recognizes the emotion "happy." This information is sent to the server, and the generating AI model recommends the most suitable content based on the following prompt message.
[2051] Example of a prompt:
[2052] "Please recommend videos that best capture the 'happiness' that users are experiencing."
[2053] This recommendation system enables personalized content delivery based on the user's emotional state, resulting in a more comfortable user experience.
[2054] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[2055] Step 1:
[2056] The user operates a device such as a smartphone to enter a request (e.g., video recommendations tailored to their emotions). The request data is entered into the device and sent to the next step.
[2057] Step 2:
[2058] The terminal device captures the user's facial image and voice data using a camera and microphone. The captured facial image and voice data become input data for analysis by the emotion engine. Digital image processing and voice analysis are performed here.
[2059] Step 3:
[2060] The terminal device sends request data and facial image / audio data to the server device. The server device receives this data and performs emotion analysis using an emotion engine. The input data is the request and emotion analysis data, and the output data is the user's emotional state (e.g., "happy").
[2061] Step 4:
[2062] The server uses an emotion engine to analyze facial images and audio data to identify the user's emotional state. Specifically, it uses facial recognition software such as "DeepFace" to determine the emotional state. The data processing performed in this process involves emotion estimation through image and audio analysis.
[2063] Step 5:
[2064] Based on the emotion engine's output data (the user's emotional state), the server uses generative AI technology to generate optimal content recommendations. Specifically, it uses the language model "GPT-3" to generate prompt sentences that correspond to the user's emotions and then makes content recommendations based on those prompts. This involves both prompt sentence generation and the generation of recommendation data using a large-scale language model.
[2065] Step 6:
[2066] The server sends content recommendation data generated by generative AI technology to the terminal device. The terminal device receives the recommendation data and displays it to the user through its interface. The data processing performed here is UI rendering, which takes the recommendation data from the server and displays it in an easy-to-understand manner for the user.
[2067] Step 7:
[2068] The user reviews the displayed content recommendations and selects the next action (e.g., playing a video). This provides feedback that improves the user experience and leads to further requests. In this step, a loop of data collection and feedback based on user interaction is crucial.
[2069] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[2070] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2071] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[2072] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2073] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[2074] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[2075] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[2076] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[2077] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[2078] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[2079] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[2080] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[2081] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[2082] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2083] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[2084] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[2085] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[2086] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[2087] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[2088] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[2089] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[2090] The following is further disclosed regarding the embodiments described above.
[2091] (Claim 1)
[2092] A terminal means for receiving requests from users,
[2093] A server means that receives a request from the terminal means and provides data generated based on the request,
[2094] A terminal means that displays the generated data received from the server means and provides an interface for the user to perform the following operations,
[2095] A system that includes this.
[2096] (Claim 2)
[2097] The system according to claim 1, characterized in that when the terminal device suggests a cooking recipe, it generates an optimal recipe based on the user's preferences and the ingredients available.
[2098] (Claim 3)
[2099] The system according to claim 1, characterized in that the terminal device generates conversation scenarios using generative AI to support English conversation practice and provides feedback to the user.
[2100] "Example 1"
[2101] (Claim 1)
[2102] A mobile device for receiving requests from users,
[2103] A computer server means that receives a request from the mobile terminal means and provides data generated based on the request,
[2104] A mobile terminal means that displays the generated data received from the computer server means and provides an interface for the user to perform the following operations,
[2105] A computing server means that generates optimal data based on requests using a generative AI model,
[2106] A mobile terminal means that displays specific instructions for the user to perform the next action based on the generated data,
[2107] A system that includes this.
[2108] (Claim 2)
[2109] The system according to claim 1, characterized in that when a mobile terminal suggests a cooking recipe, it includes a computer server that generates an optimal recipe based on the user's preferences and the ingredients available at hand.
[2110] (Claim 3)
[2111] The system according to claim 1, characterized in that it includes a computer server that generates conversation scenarios using a generative AI model and provides feedback to the user in order to support the mobile terminal means in practicing English conversation.
[2112] "Application Example 1"
[2113] (Claim 1)
[2114] A terminal device means for receiving requests from users,
[2115] A server device means that receives a request from the terminal device means and provides data generated based on the request,
[2116] A terminal device means that displays the generated data received from the server device means and provides an interface for the user to perform the following operations,
[2117] A system that includes this.
[2118] (Claim 2)
[2119] The system according to claim 1, characterized in that when the terminal device suggests a cooking recipe, it generates an optimal recipe based on the user's preferences and the ingredients available at hand.
[2120] (Claim 3)
[2121] The system according to claim 1, characterized in that the terminal device generates conversation scenarios using generative AI technology to support English conversation practice and provides feedback to the user.
[2122] (Claim 4)
[2123] The system according to claim 1, characterized in that the terminal device means generates appropriate data using generation AI technology in order to provide product descriptions, coordination suggestions, and sales information in a physical store, and provides feedback to the user and store staff.
[2124] "Example 2 of combining an emotion engine"
[2125] (Claim 1)
[2126] Information terminal device means for receiving requests from users,
[2127] A computer means that receives a request from the aforementioned information terminal device means and provides data generated based on the request,
[2128] Information terminal device means that displays the generated data received from the aforementioned computer means and provides an interface for the user to perform the following operations,
[2129] A means of emotion analysis that recognizes the user's emotional state,
[2130] A generative AI means for generating conversation scenarios and suggestions based on emotional states,
[2131] A system that includes this.
[2132] (Claim 2)
[2133] The system according to claim 1, characterized in that when the terminal information device means suggests a cooking recipe, it generates an optimal recipe based on the user's preferences and available ingredients, and further adjusts it according to the user's emotional state.
[2134] (Claim 3)
[2135] The system according to claim 1, characterized in that the terminal information device means generates conversation scenarios using a generation AI means to support English conversation practice, and the scenarios are adjusted based on the user's emotional state.
[2136] "Application example 2 when combining with an emotional engine"
[2137] (Claim 1)
[2138] A terminal means for receiving requests from users,
[2139] A server means that receives a request from the terminal means and provides data generated based on the request,
[2140] A terminal means that displays the generated data received from the server means and provides an interface for the user to perform the following operations,
[2141] An emotion engine that analyzes the user's facial image and voice data to recognize their emotional state,
[2142] A generative AI technology that generates optimal content recommendations based on the output of the aforementioned emotion engine,
[2143] A system that includes this.
[2144] (Claim 2)
[2145] The system according to claim 1, characterized in that when the terminal device suggests a cooking recipe, it generates an optimal recipe based on the user's preferences and the ingredients available.
[2146] (Claim 3)
[2147] The system according to claim 1, characterized in that the terminal device generates conversation scenarios using generative AI to support English conversation practice and provides feedback to the user. [Explanation of symbols]
[2148] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A terminal means for receiving requests from users, A server means that receives a request from the terminal means and provides data generated based on the request, A terminal means that displays the generated data received from the server means and provides an interface for the user to perform the following operations, A system that includes this.
2. The system according to claim 1, characterized in that when the terminal device suggests a cooking recipe, it generates an optimal recipe based on the user's preferences and the ingredients available.
3. The system according to claim 1, characterized in that the terminal device generates conversation scenarios using generative AI to support English conversation practice and provides feedback to the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A