System
The platform addresses inefficiencies in managing multiple generative AIs by converting prompts, integrating results, and enabling call-and-response, resulting in high-quality outputs.
Patent Information
- Application Number
- JP2024119040
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional generative AI systems lack efficient means to integrate and manage multiple generative AIs, leading to tedious task management and suboptimal quality of results due to the absence of effective call-and-response mechanisms.
A platform that allows users to input prompts, select multiple generation AIs, and includes a server to convert prompts, manage call-and-response between AIs, integrate products, and recommend AIs based on ratings and reviews.
Enables users to efficiently obtain high-quality products by seamlessly integrating multiple generative AIs and improving result quality through centralized management and collaboration.
Smart Images

Figure 2026017979000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional generative AI systems lack a means to effectively integrate and manage individual generative AIs specialized for specific tasks. As a result, when users use multiple generative AIs, tasks such as entering prompts and integrating results become tedious, requiring significant effort to obtain optimal results. Another issue is the lack of functionality for improving the quality of results through call-and-response between generative AIs. The present invention aims to solve these issues and provide a platform that allows users to efficiently obtain high-quality results. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing the following means: a means for a user to input a prompt and select multiple generation AIs, and a means for a server to convert the prompt received from the user into a format suitable for each generation AI and send it to the multiple generation AIs. Furthermore, a system is constructed that includes a means for the generation AI to generate a product based on the prompt, and a means for the server to receive and integrate the products from each generation AI.
[0006] In addition, the server manages the call and response between the generation AIs and includes a means for retransmitting the results from the first-stage generation AI to the next-stage generation AI, thereby improving the quality of the products.Furthermore, it includes a means for users to search for, select, purchase, and review new generation AIs through a generation AI app store, and a means for the server to recommend generation AIs based on ratings and reviews.These means allow users to efficiently obtain high-quality products.
[0007] "User interface" is a general term for the screens, input devices, and software components that users use to operate the platform.
[0008] A "prompt" refers to the text or data that a user uses to input instructions or information into the generated AI.
[0009] "Generative AI" refers to an artificial intelligence program that automatically creates text, images, sounds, or other forms of artifacts based on input prompts.
[0010] "Server" refers to the computer system that processes user requests and mediates interactions with the generative AI.
[0011] "Middleware" is intermediate software that receives user prompts and sends them to the appropriate generative AI.
[0012] "Generative AI App Store" refers to an online platform where users can search, select, purchase, and review various generative AIs.
[0013] "Call and response" refers to the process by which generative AIs send prompts to each other to improve the quality of their output.
[0014] "Product" is a general term for the output results (text, images, audio, etc.) created by the generative AI based on the prompts.
[0015] "Integration" refers to the process of combining and unifying the products obtained from multiple generative AIs.
[0016] "Ratings and reviews" refer to feedback and comments provided by users after using generative AI.
[0017] "Recommendation" refers to the function in which the server recommends the most suitable generation AI based on the user's past usage history and ratings. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention is a system that provides a platform for users to efficiently use multiple generation AIs to obtain high-quality products. The operation of the system that embodies the present invention will be described below.
[0040] User Interface
[0041] When users access the platform from their device, they are presented with a screen where they can select a specific project and enter prompts. In this interface, users can select each generated AI.
[0042] example:
[0043] The user inputs into the terminal, "I want to create the first act of a drama script," and selects the "drama script generation AI." He also selects the "landscape drawing AI" to draw the scenery.
[0044] Sending prompts and selecting generation AI
[0045] The terminal sends the prompts entered by the user and the information of the selected generated AI to the server, which then assigns each generated AI to the task specified by the user.
[0046] example:
[0047] The user sends a prompt, "Draw a scene where the main character uses magic for the first time," and it arrives at the server.
[0048] Prompt conversion and generation AI transmission
[0049] The server converts the received prompt into a format suitable for each generation AI and sends it to each generation AI, so that each generation AI can accurately understand and process the prompt.
[0050] example:
[0051] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," and sends it to the drama script generation AI.
[0052] Response of the generating AI and reception of the product
[0053] The server receives responses from each AI generator. The AI generator creates a product based on each prompt and returns it to the server. The server then integrates these products and provides them to the user in the most optimal form.
[0054] example:
[0055] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[0056] Call and response between generative AI
[0057] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the results. This call-and-response process enables collaboration between the generation AIs.
[0058] example:
[0059] The results of the drama script generation AI are sent to the scenery drawing AI, which then generates a new product based on that scene.
[0060] Integration and delivery of the final product
[0061] The server integrates the results from multiple generative AIs and provides the final result to the user, allowing the user to receive high-quality results all at once.
[0062] example:
[0063] The server sends the integrated results to the user as a "detailed script for the scene in which the protagonist uses magic in the square" and a "scenery description of that scene."
[0064] Use of the Generative AI App Store
[0065] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs based on the user's needs.
[0066] example:
[0067] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[0068] In this way, the present invention provides a centralized management platform that allows users to efficiently utilize multiple generation AIs and obtain optimal products.
[0069] The processing flow will be explained below.
[0070] Step 1:
[0071] The user accesses the platform from their device and logs in. This authenticates the user.
[0072] Specific behavior:
[0073] The user enters their ID and password on the login screen.
[0074] The device sends the login information to the server.
[0075] The server refers to the user database and performs authentication.
[0076] Step 2:
[0077] The user selects a project, fills in the prompts, and selects each generated AI.
[0078] Specific behavior:
[0079] On the device interface screen, select the desired project from the project list.
[0080] The user enters a prompt and selects the generated AI to use from a drop-down menu or list.
[0081] The device sends the selections and inputs to the server.
[0082] Step 3:
[0083] The server processes the received prompts and generated AI selection information and converts them into a format suitable for each generated AI.
[0084] Specific behavior:
[0085] The server parses the prompt and converts it into a format to send to each generated AI.
[0086] The server creates the necessary requests according to the generated AI's API endpoint.
[0087] Step 4:
[0088] The server sends the converted prompt to each generated AI and waits for a response.
[0089] Specific behavior:
[0090] The server sends an HTTP request to each generated AI.
[0091] Each generation AI receives a prompt and creates a product based on it.
[0092] The server waits for a response from the spawning AI.
[0093] Step 5:
[0094] After the products are returned from each generation AI, the server receives them, integrates them, and processes them.
[0095] Specific behavior:
[0096] The server receives the HTTP response from the generated AI.
[0097] Analyzes the received artifacts and prepares them for integration.
[0098] If necessary, call and response between products will be carried out.
[0099] Step 6:
[0100] When there is a call and response between generation AIs, the server sends the results of the first generation to the next generation AI.
[0101] Specific behavior:
[0102] The server sends the result from the first stage generation AI to the next stage generation AI as a prompt.
[0103] The next generation AI creates a new product and returns it to the server.
[0104] The server repeats this process as necessary.
[0105] Step 7:
[0106] The server then sends the final integrated product to the user's terminal.
[0107] Specific behavior:
[0108] The server converts the aggregated product into the appropriate format (text, images, other media).
[0109] The server sends the result to the user's terminal.
[0110] The user checks the generated data on the terminal and downloads it if necessary.
[0111] Step 8:
[0112] Users use the Generative AI App Store to search, select, purchase, and review new Generative AI.
[0113] Specific behavior:
[0114] Access the generated AI app store from your device.
[0115] Users use the search function to find generative AI.
[0116] The server displays search results, ratings, and reviews.
[0117] Users select the generated AI and make purchases or reviews.
[0118] The server recommends generative AI based on user ratings.
[0119] Through the above steps, the system allows users to efficiently obtain high-quality products.
[0120] Example 1
[0121] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0122] In today's digital service delivery environment, it is difficult for users to centrally manage a wide variety of generative AIs and obtain high-quality results. Furthermore, there is a lack of mechanisms for automatically coordinating work between different generative AIs, which often results in a lack of consistency in the quality of the results. This requires users to manually switch between individual generative AIs, which increases the complexity of the operation. Furthermore, there is an insufficient process for sharing results between generative AIs, making it difficult to improve the quality of the results.
[0123] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0124] In this invention, the server includes: a means for a user to input a prompt and select multiple generation AIs; a means for a terminal to transmit information about the prompt received from the user and the selected generation AI to the server; a means for the server to convert the prompt received from the user into a format suitable for each generation AI and transmit it to the multiple generation AIs; a means for the generation AI to generate a product based on the prompt; a means for the server to receive and integrate the products from each generation AI; and a means for the server to transmit the integrated product to the user's terminal. This allows users to efficiently use multiple generation AIs and obtain high-quality, consistent products. The server also includes a means for managing call and response between the generation AIs and retransmitting results from the first-stage generation AI to the second-stage generation AI. This process further improves the quality of the products. Furthermore, the system includes a means for a user to search for, select, purchase, and review new generation AIs through a generation AI app store, and a means for the server to recommend generation AIs based on ratings and reviews, allowing users to easily find and use the generation AI that best suits their needs.
[0125] A "user" is an entity that uses the system to select a project, input prompts, and select from multiple generation AIs.
[0126] A "prompt" is a text message that a user enters to instruct the generated AI on a specific task.
[0127] "Generative AI" is a type of artificial intelligence that generates specific outcomes based on prompts provided by the user.
[0128] A "terminal" is an electronic device that can be accessed and operated by a user and is used to input and send prompts to a server.
[0129] The "server" is a central processing unit that converts prompts received from users into a format suitable for each generation AI, sends them to the generation AI, and integrates the results to provide them to the user.
[0130] "Products" are the output results that the generative AI creates based on prompts, including text, images, and other digital content.
[0131] "Call and response" is a process of passing results between generation AIs, and is a method of improving the quality of the product by retransmitting the results obtained from the first-stage generation AI to the next-stage generation AI.
[0132] "Generative AI App Store" means an online platform where users can search, select, purchase, and review generative AI.
[0133] "Recommendation" is the act of the server selecting and suggesting an appropriate generation AI based on the user's needs, ratings, and reviews.
[0134] The present invention provides a platform that allows users to efficiently use multiple generation AIs to obtain high-quality products. An embodiment of this system will be described in detail below.
[0135] User Interface
[0136] Users access the platform through a web browser or dedicated application using an internet-connected device (e.g., personal computer, tablet, smartphone, etc.). In this interface, users are presented with a screen to select a specific project and a screen to enter prompts. After selecting a project and entering prompts, users can select the generative AI to use.
[0137] Specific examples
[0138] The user inputs "I want to create the first act of a drama script" and selects "Drama script generation AI." They can also select "Landscape drawing AI" for drawing scenery.
[0139] Prompt and generate AI selection information
[0140] After the user selects a prompt and a generated AI, the device sends this information to the server in a structured data format such as JSON.
[0141] Prompt Translation
[0142] The server converts the prompts received from the device into a format that is easy for the AI to understand. This process is performed using a natural language processing module. The server converts the prompts into a format suitable for each AI and sends them to each AI.
[0143] Specific examples
[0144] The server converts the prompt "Please describe the scene where the main character uses magic for the first time" into "Scene 1: The main character uses magic for the first time in the square. Describe the scene and emotions in detail," and sends this to the drama script generation AI.
[0145] Receiving the generated AI's response
[0146] The generation AI creates a product based on the prompts and sends the results back to the server. The server receives these responses and integrates them. Specifically, it combines the results of the drama script generation AI with the results of the scenery drawing AI.
[0147] Specific examples
[0148] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[0149] Improving product quality through call and response
[0150] The server improves the quality of the generated results by retransmitting the results of the first-stage generation AI to the next-stage generation AI. This process is called call and response, and allows results to be passed between multiple generation AIs.
[0151] Specific examples
[0152] The server sends the results of the drama script generation AI to the scenery drawing AI, which then uses that information to generate new results. For example, the scenery drawing AI might return a result such as "depict the scenery at the moment the main character casts a spell in more detail."
[0153] Integration and delivery of the final product
[0154] The server aggregates the final product and sends it to the user's device, ensuring that the user receives a consistent, high-quality product at once.
[0155] Specific examples
[0156] The server sends the integrated results to the user as a "detailed script for the scene where magic is used in the square" and a "scenery description of that scene."
[0157] Use of the Generative AI App Store
[0158] Users can search, select, purchase, and review new AI generators through the AI generator app store, and the server will recommend the best AI generators to users based on these ratings and reviews.
[0159] Specific examples
[0160] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[0161] As described above, the present invention provides a unified management platform that allows users to efficiently utilize multiple generative AIs and obtain high-quality, consistent results.
[0162] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0163] Step 1:
[0164] The user selects a project and enters a prompt. This information is entered by accessing the platform from the user's device. The user selects the prompt and the generation AI to use.
[0165] Specific actions
[0166] The user accesses the platform's web page using a browser or app. On the project selection screen, they select "Drama Script" and then enter "Please create a scene where the main character uses magic for the first time" on the prompt input screen. They also select "Drama Script Generation AI" and "Landscape Drawing AI" as the generation AIs to use.
[0167] input
[0168] The prompt entered by the user, "Please create a scene where the main character uses magic for the first time," and the selection information of the generated AI.
[0169] output
[0170] The user's input data is temporarily saved in the device's internal storage.
[0171] Step 2:
[0172] The terminal transmits the prompt received from the user and information about the selected generated AI to the server.
[0173] Specific actions
[0174] When the user clicks the "Submit" button, the device converts the prompt and AI selection information into JSON format and sends it to the server.
[0175] input
[0176] User-entered prompts and generated AI selection information.
[0177] output
[0178] The data is sent to the server in JSON format.
[0179] Step 3:
[0180] The server converts the prompts received from the user into a format suitable for each generation AI and sends them to multiple generation AIs.
[0181] Specific actions
[0182] The server uses a natural language processing module to analyze the prompt and convert it into an input format for each generation AI. For example, for the "drama script generation AI," it converts it into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," and for the "scenery drawing AI," it converts it into "Describe the scene at the moment the magical light is emitted in the square."
[0183] input
[0184] The prompt and generated AI selection information received by the server in JSON format.
[0185] output
[0186] The prompt is converted into a format suitable for the generation AI and sent to the corresponding generation AI.
[0187] Step 4:
[0188] The generation AI generates artifacts based on prompts.
[0189] Specific actions
[0190] Each AI generator uses its internal algorithm to create a result based on the prompts it receives. For example, the "drama script generator AI" generates a scenario, and the "landscape drawing AI" generates an illustration.
[0191] input
[0192] The prompt converted into a format suitable for generative AI.
[0193] output
[0194] The artifacts (scripts, illustrations, etc.) are generated and sent back to the server.
[0195] Step 5:
[0196] The server receives and integrates the products from each generation AI.
[0197] Specific actions
[0198] The server receives the response data from each generation AI and combines them into a single integrated product, for example, integrating a scene from a drama script with the corresponding landscape description.
[0199] input
[0200] The products received from each generating AI.
[0201] output
[0202] Integrated artifacts (e.g., "A detailed script for the magic scene in the square" and "A landscape description of the scene").
[0203] Step 6:
[0204] The server manages the call and response between the generation AIs and retransmits the results from the first generation AI to the next generation AI in order to improve the quality of the results.
[0205] Specific actions
[0206] The server retransmits the results of the drama script generation AI to the scenery drawing AI, which then generates new products based on that information.
[0207] input
[0208] Results generated by the first-stage generation AI.
[0209] output
[0210] Improved production from next-stage generation AI.
[0211] Step 7:
[0212] The server aggregates the final product and sends it to the user's device.
[0213] Specific actions
[0214] The server integrates the products obtained from all the generation AIs, formats them as the final product, and provides it to the user.
[0215] input
[0216] A product integrated from multiple generation AIs.
[0217] output
[0218] The final product is sent to the user's terminal.
[0219] Step 8:
[0220] Users search, select, purchase, and review new generative AI through the generative AI app store.
[0221] Specific actions
[0222] Users access the app store, search for and purchase a generative AI, and write a review. The server analyzes these ratings and reviews and recommends the best generative AI for the user.
[0223] input
[0224] User ratings and reviews.
[0225] output
[0226] Recommendation information from the server.
[0227] (Application example 1)
[0228] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0229] In conventional design generation systems using generative AI models, it was difficult to effectively integrate multiple generative AI models and easily create high-quality 3D designs and product introduction videos for virtual stores. Furthermore, collaboration between generative AI models was insufficient, preventing sufficient improvement in the quality of the results. Furthermore, users lacked appropriate information when searching for and selecting new generative AI models, making it difficult to find the optimal model.
[0230] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0231] In this invention, the server includes means for a user to input a prompt sentence and select multiple generative AI models, means for the server to convert the prompt sentence received from the user into a format suitable for each generative AI model and send it to the multiple generative AI models, means for the generative AI models to generate products based on the prompt sentence, means for the server to receive and integrate the products from each generative AI model, and means for the operator of the virtual store to generate high-quality 3D designs and product introduction videos. This allows users to effectively use multiple generative AI models and obtain high-quality products at once, making it possible to generate high-quality content in the virtual store.
[0232] "User device" means an electronic device operated by a user, including a smartphone, smart glasses, a head-mounted display, etc.
[0233] A "generative AI model" refers to an artificial intelligence algorithm that automatically creates a specific product based on a prompt entered by a user.
[0234] A "prompt" is a text expression that succinctly summarizes the input information and instructions that a user provides to a generative AI model.
[0235] "Products" are the deliverables that the generative AI model creates based on the prompt, including 3D designs and product introduction videos.
[0236] A "server" is a central processing unit that runs on a network and manages the exchange of data between the user's device and the generative AI model.
[0237] A "virtual store" is a virtual store that displays and sells products and services online.
[0238] "Call and response between generative AI models" is a technique in which generative AI models work together to repeatedly respond, improving the accuracy and quality of the products.
[0239] A "Generative AI Model App Store" is an online platform where users can search for, select, purchase, and review new generative AI models.
[0240] A system for implementing the present invention includes a user terminal, a server, and multiple generative AI models. The user terminal is an electronic device such as a smartphone, smart glasses, or a head-mounted display, and provides an interface for a user to input a prompt sentence and select a generative AI model.
[0241] Hardware and software used
[0242] Hardware: Smartphones (e.g., iPhone, Samsung Galaxy), smart glasses (e.g., Google Glass), head-mounted displays (e.g., Meta Quest)
[0243] Software: iOS / Android app, web server, generative AI model (e.g., OpenAI's ChatGPT, DALL-E)
[0244] System Operation
[0245] 1. User's device:
[0246] Users access the application from their device, select a specific project, enter a prompt, and can select the generative AI model they want to use.
[0247] example:
[0248] I want to create a 3D model of a new product, and I need high-resolution images as well as a promotional video.
[0249] At this time, the user can select product design AI, video generation AI, or landscape drawing AI.
[0250] 2. Server prompt translation:
[0251] The server converts the prompt received from the user into a format suitable for each generative AI model, and then sends the converted prompt to each generative AI model.
[0252] 3. How generative AI models work:
[0253] The generative AI model creates a specified product based on the received prompt. For example, a product design AI generates a 3D model, a video generation AI generates a product introduction video, and a landscape drawing AI generates a background design.
[0254] 4. Server integration process:
[0255] The server receives the results from each generative AI model and integrates them. The integrated results are then sent to the user's device. At this time, a call-and-response function between the generative AI models is also activated, and the results of multiple generative AI models are sequentially linked to improve the quality of the results.
[0256] 5. User Interface:
[0257] The user can again review the product through the application, provide feedback on the product if necessary, and enter further prompts.
[0258] 6. Generative AI model app store:
[0259] Users can use the generative AI model app store to search for, select, purchase, and review new generative AI models, and the server has the function of recommending generative AI models based on ratings and reviews.
[0260] In this way, users can efficiently utilize a variety of generative AI models to easily generate high-quality 3D designs and product introduction videos for virtual stores.
[0261] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0262] Step 1:
[0263] A user accesses the application from their device, selects a specific project, and enters a prompt. The user selects the generative AI model to use, such as product design AI, video generation AI, or landscape drawing AI. The prompt entered might be something like, "I would like to create a 3D model of a new product. I also need a promotional video along with high-resolution images."
[0264] Step 2:
[0265] The terminal transmits the input prompt sentence and information about the selected generative AI model to the server. At this time, the input data from the terminal includes the prompt sentence and the ID information of the generative AI model.
[0266] Step 3:
[0267] The server analyzes the prompt received from the user and converts it into a format suitable for each generative AI model. For example, a prompt such as "I would like to create a 3D model of a new product" is converted into a format suitable for a product design AI to generate a high-resolution 3D model, and the part "We also need an introductory video" is converted into a format suitable for a video generation AI. The converted prompt is then sent to the generative AI model.
[0268] Step 4:
[0269] The generative AI models (product design AI, video generation AI, and landscape drawing AI) create the specified product based on the received prompt. The product design AI generates a high-resolution 3D model, the video generation AI generates a video introducing the product, and the landscape drawing AI generates a background design. Each generative AI model uses its internal algorithm to calculate data based on the prompt.
[0270] Step 5:
[0271] Each generative AI model completes the product and returns the data to the server. For example, a product design AI sends 3D model data, a video generation AI sends a video file, and a landscape drawing AI sends a background image.
[0272] Step 6:
[0273] The server integrates the results received from each generative AI model. It uses an integration algorithm to optimally combine the data, combining the 3D model data, video files, and background images into a single piece of content. During this integration process, a call-and-response function between the generative AI models is utilized to adjust the results to ensure higher quality.
[0274] Step 7:
[0275] The server transmits the integrated product to the user's terminal, which displays the received product and an interface for providing feedback if necessary.
[0276] Step 8:
[0277] When a user uses the generative AI model app store, the user can search, select, purchase, and review new generative AI models, and the server recommends the most suitable generative AI models based on the user's ratings and reviews.
[0278] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0279] The present invention provides a system that allows users to efficiently use multiple AI generators and further improves the quality of the output by recognizing the user's emotional state. The following describes in detail how the present invention is implemented.
[0280] User Interface
[0281] The process begins when a user accesses the platform from their device and logs in. The user selects a project on the interface screen, fills in prompts, and selects the generative AI to use. This interface works in conjunction with an emotion engine that also recognizes the user's emotions.
[0282] example:
[0283] The user inputs "I want to create the first act of a drama script" into the device and selects the "drama script generation AI." They also select the "landscape drawing AI" to draw the scenery. The emotion engine then analyzes the user's voice and facial expressions and recognizes them as "excited."
[0284] Sending prompts and selecting generation AI
[0285] The device sends the prompts entered by the user, information on the selected generation AI, and emotional data obtained from the emotion engine to the server, which then adjusts the prompts according to the user's emotional state.
[0286] example:
[0287] The prompt sent by the user, "Draw a scene where the main character uses magic for the first time," is received by the server, and the emotion engine detects the "excited" state, so the prompt is adjusted slightly.
[0288] Prompt conversion and generation AI transmission
[0289] The server converts the received prompts and emotion data into a format suitable for each generation AI and sends it to each generation AI. Data obtained from the emotion engine is also applied, so the generated content is adjusted to match the user's emotion.
[0290] example:
[0291] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," adds emotional data, and sends it to the drama script generation AI.
[0292] Response of the generating AI and reception of the product
[0293] The server receives responses from each AI generator. The AI generator creates a product based on each prompt and returns it to the server. The server then integrates these products and provides them to the user in the most optimal form.
[0294] example:
[0295] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[0296] Call and response between generative AI
[0297] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the results. This call-and-response process allows the generation AIs to cooperate with each other. Data from the emotion engine is also continuously used.
[0298] example:
[0299] The results of the drama script generation AI are sent to the scenery drawing AI, which then generates a new product based on that scene.
[0300] Integration and delivery of the final product
[0301] The server integrates the results from multiple generative AIs and provides the final result to the user, allowing users to receive high-quality results at once. The result is also further adjusted based on emotional data.
[0302] example:
[0303] The server sends the integrated results to the user as a detailed script for the scene in which the protagonist uses magic in the square and a description of the scenery of that scene. The description of the scene is adjusted to be even more dramatic depending on the user's level of excitement.
[0304] Use of the Generative AI App Store
[0305] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs based on the user's needs. This process also utilizes data from the emotion engine to recommend Generative AIs that are tailored to the user's current emotional state.
[0306] example:
[0307] Users select and install a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project." Users who are in an "excited" state will be suggested generative AIs that will further increase their excitement.
[0308] In this way, the present invention provides a centralized management platform for users to efficiently utilize multiple generative AIs and further improve the quality of the output using an emotion engine.
[0309] The processing flow will be explained below.
[0310] Step 1:
[0311] The user accesses the platform from their device and logs in. Login authentication is performed and the user's personal information is confirmed.
[0312] Specific behavior:
[0313] The user enters their ID and password on the login screen.
[0314] The device sends the login information to the server.
[0315] The server refers to the user database and performs authentication.
[0316] If the authentication is successful, the user will be taken to the dashboard screen.
[0317] Step 2:
[0318] The user selects a project, enters prompts, and selects each generated AI. The emotion engine then analyzes the user's emotional state.
[0319] Specific behavior:
[0320] On the device interface screen, select the desired project from the project list.
[0321] The user enters a prompt and selects the generated AI to use from a drop-down menu or list.
[0322] An emotion engine built into or connected to the device collects and analyzes emotional data from the user's voice and facial expressions.
[0323] The device sends prompts, the generated AI's selection information, and emotion data to the server.
[0324] Step 3:
[0325] The server receives the sent prompt, the selected generation AI information, and the emotion data from the emotion engine, and converts the prompt into a format suitable for each generation AI based on that information.
[0326] Specific behavior:
[0327] The server parses the prompt and converts it into a format to send to the generating AI.
[0328] Prepare to use emotion data obtained from the emotion engine to adjust the content of prompts and productions.
[0329] Create the necessary requests according to the API endpoint of each generation AI.
[0330] Step 4:
[0331] The server sends the converted prompt to each generation AI, requesting it to generate a product.
[0332] Specific behavior:
[0333] The server sends an HTTP request to the generated AI's API endpoint.
[0334] Each generation AI receives a prompt and creates the specified product.
[0335] Step 5:
[0336] Each generation AI creates a product and returns a response to the server, which receives these products.
[0337] Specific behavior:
[0338] Each generation AI generates a product based on a prompt.
[0339] The server receives the HTTP response from the generation AI and obtains the generated data.
[0340] Step 6:
[0341] The products received by the server are further adjusted and integrated through call and response between each generation AI.
[0342] Specific behavior:
[0343] The server sends the result from the first stage generation AI to the next stage generation AI as a prompt.
[0344] The next generation AI creates a new product and returns it to the server.
[0345] Repeat this process as necessary.
[0346] At each stage, the emotion engine data is used to adjust the output.
[0347] Step 7:
[0348] The final integrated product is sent to the user's device by the server, where it is further adjusted based on the emotion data.
[0349] Specific behavior:
[0350] The server converts the aggregated product into the appropriate format (text, images, other media).
[0351] The server further adjusts the product based on data from the emotion engine and sends it to the user's device.
[0352] Step 8:
[0353] Users can review the results they receive, make corrections and provide feedback as needed, and use the Generative AI App Store to find new Generative AI.
[0354] Specific behavior:
[0355] The user checks the output on the terminal.
[0356] Make corrections as needed and send feedback to the server.
[0357] The device accesses the generative AI app store, and the user searches for, selects, and purchases new generative AI.
[0358] The server recommends a generative AI based on user feedback.
[0359] Example 2
[0360] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0361] Conventional generative AI systems make it difficult for users to efficiently use multiple generative AIs, and have limited means for improving the consistency and quality of the results. Furthermore, they often fail to meet user expectations because they are unable to adjust the results to take into account the user's emotional state. The present invention aims to solve these problems and provide high-quality results that respond to the user's emotional state.
[0362] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for converting prompts and emotion data received from the user into a format suitable for each generation AI and transmitting the converted data to the multiple generation AIs, means for the generation AIs to generate products based on the prompts and emotion data, and means for the server to receive and integrate the products from the generation AIs. This makes it possible to adjust prompts based on the user's emotional state and consistently provide high-quality products.
[0363] A "user" is a person or institution that utilizes the system to enter prompts and select a generating AI.
[0364] "Terminal" means an electronic device used by a user to access the system and provide input information.
[0365] The "server" is a central processing unit that processes information received from users, converts it into a format suitable for the generative AI model, integrates the generated results, and provides them to users.
[0366] "Generative AI" is an artificial intelligence model that creates a specified artifact based on user-entered prompts.
[0367] A "prompt" is an instruction that the user enters to specify the desired product for the generation AI.
[0368] A "product" is an output result that the generation AI generates based on a prompt.
[0369] "Emotional data" refers to data that indicates the user's current emotional state and is used to adjust the product.
[0370] The "emotion engine" is a function that analyzes the user's voice, facial expressions, etc. to generate emotional data.
[0371] "Call and response" is a process in which the results obtained from the first-stage generation AI are retransmitted to the next-stage generation AI to improve the quality of the product.
[0372] The "Generative AI App Store" is an online platform where users can search for, select, purchase, and review new generative AI.
[0373] "Recommendation" means that the server suggests a generative AI that is suitable for the user based on ratings and reviews.
[0374] "Format conversion" is the process by which the server converts the prompt and emotion data received from the user into a format suitable for the generative AI.
[0375] "Integration" refers to the server combining the products obtained from each generation AI into one.
[0376] The system of the present invention is a platform that allows users to efficiently utilize multiple generation AIs and improve the quality of the products based on the user's emotional state. The components and specific functions of this system will be described below.
[0377] User Interface
[0378] Users access the platform from their own device and log in. Next, they select a project and enter a prompt for the AI generator. The user then selects the AI generator they want to use. At this time, an emotion engine is connected to the device, which analyzes and recognizes the user's emotional state in real time.
[0379] Example prompt sentence:
[0380] "I want to create the first act of a drama script."
[0381] Select "Drama script generation AI" and also select "Landscape drawing AI" for landscape drawing.
[0382] The emotion engine analyzes the user's voice and facial expressions to detect an "excited" state.
[0383] Sending prompt information and emotion data
[0384] The device sends the prompt text entered by the user, the selected generation AI information, and the emotion data obtained from the emotion engine to the server, which receives this information and prepares it for processing.
[0385] Prompt adjustment and transformation
[0386] The server analyzes the received prompt and emotion data, adjusts the prompt content as needed, and then converts it into the optimal format for each generative AI model before sending it.
[0387] Examples:
[0388] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," adds emotional data about excitement, and sends it to the drama script generation AI.
[0389] Receiving and integrating generative AI responses
[0390] The server receives responses from each AI generator. Each AI generator creates a product based on the prompt and emotion data and sends it back to the server. The server then integrates these products and provides them to the user in the most optimal form.
[0391] Examples:
[0392] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[0393] Call and response between generative AI
[0394] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the product. In this process, collaboration between the generation AIs is realized, and emotion data is also continuously used.
[0395] Examples:
[0396] The server sends the results of the drama script generation AI to the scenery drawing AI, which then generates a new product based on that scene.
[0397] Providing the final product
[0398] The server integrates the results from multiple generative AIs and provides the final result to the user. The result is further adjusted based on the user's emotional data, allowing the user to receive high-quality results all at once.
[0399] Examples:
[0400] The server sends the integrated results to the user as a "detailed script for the scene in which the protagonist uses magic in the square" and a "scenery description of that scene."
[0401] The depiction of the scene is adjusted to become more dramatic depending on the user's level of excitement.
[0402] Use of the Generative AI App Store
[0403] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs that meet the user's needs. Data from the emotion engine is also utilized to suggest Generative AIs that are suited to the user's current emotional state.
[0404] Examples:
[0405] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[0406] For users who are excited, generative AI will be suggested to further increase their excitement.
[0407] In this way, the present invention provides a unified platform that allows users to efficiently utilize multiple generative AI models and provide the highest quality products based on emotion data.
[0408] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0409] Step 1:
[0410] A user accesses the platform from their device and logs in. At this time, the user enters authentication information (username, password) and is granted access to the system. If the entered authentication information is correct, the server starts a session and displays the user interface (UI). This UI includes a project selection screen and a generation AI selection screen.
[0411] Input: Username, Password
[0412] Output: Session starts, UI is displayed
[0413] Specific behavior:
[0414] The user enters their credentials on the login page and clicks the submit button.
[0415] The server verifies the authentication information and, if correct, displays the dashboard screen.
[0416] Step 2:
[0417] The user selects a project on the interface screen and inputs a prompt for the AI to generate. The user also selects the AI to use. The device is connected to an emotion engine that analyzes and recognizes the user's emotional state in real time.
[0418] Input: Project selection, prompt text, AI generation selection
[0419] Output: Project information, prompt information, emotion data
[0420] Specific behavior:
[0421] The user enters "I want to create the first act of a drama script" and selects "Drama script generation AI."
[0422] The emotion engine scans the user's facial expressions to detect "excitement."
[0423] Step 3:
[0424] The device sends the prompt text entered by the user, information on the selected generation AI, and emotion data obtained from the emotion engine to the server, which receives this information and prepares to adjust the prompt text.
[0425] Input: prompt, generated AI information, emotion data
[0426] Output: Send data to the server
[0427] Specific behavior:
[0428] The terminal sends project information, prompts and "excitement state" data to the server.
[0429] Step 4:
[0430] The server analyzes the received prompt and emotion data, adjusts the prompt as needed, and then converts it into a format optimal for each generative AI model and sends it.
[0431] Input: prompt sentence, emotion data
[0432] Output: Adjusted prompt sentence, data sent to the generation AI
[0433] Specific behavior:
[0434] The server translates the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail."
[0435] Data adjusted to a format suitable for the generative AI model is sent to each generative AI.
[0436] Step 5:
[0437] The generation AI generates a product based on the prompt sentence and emotion data and sends it back to the server. The server receives responses from each generation AI and integrates the products.
[0438] Input: Adjusted prompt sentence, emotion data
[0439] Output: Artifacts (e.g. text, images)
[0440] Specific behavior:
[0441] Generative AI generates text and images based on the prompts it receives.
[0442] The server receives and integrates the products returned by the generation AI.
[0443] example:
[0444] Response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience erupts in amazement and cheers."
[0445] The landscape drawing AI responds, "The square is lush and green, and the magic light is blue."
[0446] Step 6:
[0447] The server retransmits the results of the first-stage AI generation to the second-stage AI generation to further improve the quality of the product. Emotion data is also continuously used. This results in a product of improved quality.
[0448] Input: Generation results from the first stage generation AI, emotion data
[0449] Output: Refined product from next stage generation AI
[0450] Specific behavior:
[0451] The server sends the results of the first stage drama script generation AI to the scenery drawing AI.
[0452] The landscape drawing AI generates new products based on the received data and sends them back.
[0453] Step 7:
[0454] The server integrates the results from multiple generative AIs and provides the final result to the user. Emotional data is also taken into account, allowing the user to receive the results in the most optimal way.
[0455] Input: Products from each generation AI, emotion data
[0456] Output: The integrated final product
[0457] Specific behavior:
[0458] The server integrates the results of each generation AI to create the final scene script and scenery description.
[0459] The scene depiction is adjusted to be even more dramatic depending on the user's level of excitement.
[0460] Step 8:
[0461] Users access the generative AI app store to search, select, purchase, and review new generative AIs. The server then recommends generative AIs that meet the user's needs, utilizing emotional data to make suggestions.
[0462] Input: User searches, selections, purchase information, ratings and reviews, sentiment data
[0463] Output: Recommendations and generative AI
[0464] Specific behavior:
[0465] A user searches for a new generative AI in the app store, checks reviews, and then purchases it.
[0466] The server recommends, "This generative AI is suitable for your project."
[0467] Through these specific steps and operations, the system of the present invention provides a high-performance and user-friendly generative AI platform.
[0468] (Application example 2)
[0469] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0470] Conventional generative AI systems lack the ability to customize their outputs based on the user's emotional state. This can result in outputs that do not match the user's expectations or emotional state, resulting in a poor user experience. Furthermore, there is also the issue of insufficient collaboration between generative AIs, resulting in a decline in the quality of the outputs. Furthermore, it is difficult to easily search for and select new generative AIs, preventing users from effectively utilizing a wide range of generative AIs.
[0471] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0472] In this invention, the server includes means for recognizing the emotional state of the user and adjusting the quality of the product based on the emotional data, an emotion engine for analyzing the emotional data, means for managing call and response between the generation AIs and retransmitting results from the first-stage generation AI to the next-stage generation AI to improve the quality of the product, means for the user to search, select, purchase, and review new generation AIs through a generation AI app store, and means for the server to recommend generation AI models based on the ratings and reviews. This makes it possible to provide high-quality products tailored to the user's emotional state, strengthens cooperation between generation AIs, and makes it easier to use new generation AIs.
[0473] "User" refers to the person using the generative AI system to input prompts.
[0474] A "prompt" refers to a sentence in the form of specific instructions or questions that are input to the generating AI.
[0475] "Generative AI" refers to an artificial intelligence model that creates artifacts based on prompts from the user.
[0476] "Server" refers to the computer system that receives prompts from the user, converts them into a format suitable for the generating AI, and then integrates and provides the resulting results to the user.
[0477] "Product" refers to the output data that the generative AI creates based on the prompts.
[0478] "Emotional state" refers to the user's current emotional or mental state, as recognized by the emotion engine.
[0479] "Emotion Engine" refers to a software or hardware system that analyzes a user's emotional state and provides that data.
[0480] "Call and response" refers to the collaborative process of retransmitting the product obtained from the first-stage generation AI to the next-stage generation AI to improve the quality of the product.
[0481] "Generative AI App Store" refers to an online platform where users can search for, select, purchase, and review new generative AI.
[0482] "Recommendation" refers to the act of suggesting the most suitable generative AI based on user ratings and reviews.
[0483] "User's device" refers to the keywords for smartphones, computers, tablets, etc. that users use to operate the generated AI system.
[0484] The present invention provides a system that provides high-quality products by efficiently utilizing multiple generation AIs while taking into account the emotional state of the user. Detailed embodiments of the present invention will be described below.
[0485] System Overview
[0486] The system of the present invention includes means for a user to input a prompt and select a generation AI, means for a server to convert the prompt into a suitable format and send it to the generation AI, means for the generation AI to generate a product based on the prompt, means for integrating the product and providing it to the user, and means for using an emotion engine to analyze the user's emotional state and adjust the product based thereon.
[0487] Hardware and software used
[0488] Hardware
[0489] User devices: smartphones, computers, tablets, etc.
[0490] Camera module: For capturing images to analyze the user's facial expressions
[0491] software
[0492] Emotion recognition engine: Software that analyzes the user's emotional state (e.g., the Hugging Face transformer library)
[0493] Generative AI models: Artificial intelligence models that create artifacts based on prompts (e.g., GPT-3)
[0494] Review analysis engine: Software that analyzes reviews based on artifacts
[0495] System Operation
[0496] 1. User Interface
[0497] Users access the interface from a device such as a smartphone or computer and log in. They then perform a search, select a project, fill in the prompts, and can also select a generative AI model.
[0498] 2. Acquiring Emotion Data
[0499] The camera module of the user's device is used to capture facial images, which are then analyzed by an emotion recognition engine to understand the user's emotional state, for example, determining whether the user is in an "excited" state.
[0500] 3. Sending prompts and selecting generation AI
[0501] The prompts and emotion data entered by the user are sent to the server, which converts them into a suitable format and sends them to the selected generative AI model.
[0502] 4. Product generation and preparation
[0503] The generative AI creates products based on prompts and adjusts the content based on emotional data. For example, if a user is in an excited state and enters the prompt, "I'm looking for a black leather jacket. Can you recommend any products?", the generative AI will generate product recommendations that correspond to the user's excited state.
[0504] 5. Consolidating and Presenting Results
[0505] The server integrates the results from each AI generator and provides it to the user, who then adjusts the results to match the user's emotional state.
[0506] 6. Use of Generative AI App Store
[0507] Users can search, select, purchase, and review new generative AI models through the generative AI app store, and the server will recommend the most suitable generative AI based on user ratings and reviews.
[0508] Specific examples
[0509] If a user types a prompt like "I'm looking for a black leather jacket. Can you recommend some?", and the user's emotional state is recognized as "excited," the following happens:
[0510] 1. Prompt and emotion data acquisition
[0511] Prompt: "I'm looking for a black leather jacket. Can you recommend one?"
[0512] Emotional state: "Excited"
[0513] 2. Product generation
[0514] The generative AI model generates detailed product descriptions and recommendations that match the user's excitement level, such as, "This black leather jacket is trending this season and has received very high ratings."
[0515] In this way, the present invention can provide high quality products based on the user's emotional state.
[0516] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0517] Step 1:
[0518] A user uses a terminal to access the interface and log in. The user selects a project and fills in the prompts. Then, the user selects multiple generative AI models.
[0519] Input: User login information, project selection, prompt text, and selection of generative AI model
[0520] Output: Information about the input prompt and the selected generative AI model
[0521] Specific behavior: For example, the user enters the prompt, "I'm looking for a black leather jacket. Can you recommend some?" and selects the relevant generative AI model.
[0522] Step 2:
[0523] The device's camera module is used to capture the user's facial expression, and the image data is sent to the emotion recognition engine for analysis.
[0524] Input: User's face image
[0525] Output: Parsed user emotional state data
[0526] Specific operation: A facial image is captured with a camera, sent to an emotion recognition engine, and the user's emotional state (e.g., excitement) is analyzed.
[0527] Step 3:
[0528] The prompt sentence and emotional state data are sent to the server, which receives the prompt sentence and emotional state data and converts the prompt sentence into a format suitable for each generative AI model.
[0529] Input: prompt sentence, emotional state data
[0530] Output: A prompt in a format suitable for a generative AI model
[0531] What happens: The server translates the prompt to something like, "I'm looking for a black leather jacket. Can you recommend one? [Emotion: Excited]."
[0532] Step 4:
[0533] The server sends the converted prompt sentence to each generative AI model, which creates a product based on the prompt sentence and returns it to the server.
[0534] Input: A prompt in a format suitable for a generative AI model
[0535] Output: The product of the generative AI model
[0536] Specific operation: The generative AI model generates product information and returns a result such as "This black leather jacket is a trendy item this season and has received very high ratings" to the server.
[0537] Step 5:
[0538] The server receives the outputs from each generative AI model, adjusts them based on the emotional state data, and integrates them.
[0539] Input: Product from generative AI model, emotional state data
[0540] Output: The integrated product
[0541] What it does: The server integrates the generated product details and recommendations based on them, dynamically adjusting them to fit the sentiment.
[0542] Step 6:
[0543] The server sends the integrated product to the user's terminal, and the user checks the received product.
[0544] Input: The integrated product
[0545] Output: The artifact provided to the user
[0546] Specific operation: The server sends a message to the user's device such as "Recommended product: This black leather jacket is a trendy item this season and has received very high ratings."
[0547] Step 7:
[0548] Users search, select, purchase, and review new generative AI models through the generative AI app store, and the server recommends generative AI models based on ratings and reviews.
[0549] Input: User ratings and reviews
[0550] Output: A generative AI model recommended to the user
[0551] Specific operation: The user searches for new generative AI models in the generative AI app store, purchases them, and reviews them. The server then uses this information to suggest the optimal generative AI model.
[0552] Through the above processing steps, this system realizes the operation of an advanced generative AI model that takes into account the user's emotional state, providing high-quality results.
[0553] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0554] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0555] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0556] [Second embodiment]
[0557] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0558] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0559] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0560] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0561] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0562] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0563] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0564] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0565] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0566] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0567] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0568] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0569] The present invention is a system that provides a platform for users to efficiently use multiple generation AIs to obtain high-quality products. The operation of the system that embodies the present invention will be described below.
[0570] User Interface
[0571] When users access the platform from their device, they are presented with a screen where they can select a specific project and enter prompts. In this interface, users can select each generated AI.
[0572] example:
[0573] The user inputs into the terminal, "I want to create the first act of a drama script," and selects the "drama script generation AI." He also selects the "landscape drawing AI" to draw the scenery.
[0574] Sending prompts and selecting generation AI
[0575] The terminal sends the prompts entered by the user and the information of the selected generated AI to the server, which then assigns each generated AI to the task specified by the user.
[0576] example:
[0577] The user sends a prompt, "Draw a scene where the main character uses magic for the first time," and it arrives at the server.
[0578] Prompt conversion and generation AI transmission
[0579] The server converts the received prompt into a format suitable for each generation AI and sends it to each generation AI, so that each generation AI can accurately understand and process the prompt.
[0580] example:
[0581] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," and sends it to the drama script generation AI.
[0582] Response of the generating AI and reception of the product
[0583] The server receives responses from each AI generator. The AI generator creates a product based on each prompt and returns it to the server. The server then integrates these products and provides them to the user in the most optimal form.
[0584] example:
[0585] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[0586] Call and response between generative AI
[0587] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the results. This call-and-response process enables collaboration between the generation AIs.
[0588] example:
[0589] The results of the drama script generation AI are sent to the scenery drawing AI, which then generates a new product based on that scene.
[0590] Integration and delivery of the final product
[0591] The server integrates the results from multiple generative AIs and provides the final result to the user, allowing the user to receive high-quality results all at once.
[0592] example:
[0593] The server sends the integrated results to the user as a "detailed script for the scene in which the protagonist uses magic in the square" and a "scenery description of that scene."
[0594] Use of the Generative AI App Store
[0595] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs based on the user's needs.
[0596] example:
[0597] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[0598] In this way, the present invention provides a centralized management platform that allows users to efficiently utilize multiple generation AIs and obtain optimal products.
[0599] The processing flow will be explained below.
[0600] Step 1:
[0601] The user accesses the platform from their device and logs in. This authenticates the user.
[0602] Specific behavior:
[0603] The user enters their ID and password on the login screen.
[0604] The device sends the login information to the server.
[0605] The server refers to the user database and performs authentication.
[0606] Step 2:
[0607] The user selects a project, fills in the prompts, and selects each generated AI.
[0608] Specific behavior:
[0609] On the device interface screen, select the desired project from the project list.
[0610] The user enters a prompt and selects the generated AI to use from a drop-down menu or list.
[0611] The device sends the selections and inputs to the server.
[0612] Step 3:
[0613] The server processes the received prompts and generated AI selection information and converts them into a format suitable for each generated AI.
[0614] Specific behavior:
[0615] The server parses the prompt and converts it into a format to send to each generated AI.
[0616] The server creates the necessary requests according to the generated AI's API endpoint.
[0617] Step 4:
[0618] The server sends the converted prompt to each generated AI and waits for a response.
[0619] Specific behavior:
[0620] The server sends an HTTP request to each generated AI.
[0621] Each generation AI receives a prompt and creates a product based on it.
[0622] The server waits for a response from the spawning AI.
[0623] Step 5:
[0624] After the products are returned from each generation AI, the server receives them, integrates them, and processes them.
[0625] Specific behavior:
[0626] The server receives the HTTP response from the generated AI.
[0627] Analyzes the received artifacts and prepares them for integration.
[0628] If necessary, call and response between products will be carried out.
[0629] Step 6:
[0630] When there is a call and response between generation AIs, the server sends the results of the first generation to the next generation AI.
[0631] Specific behavior:
[0632] The server sends the result from the first stage generation AI to the next stage generation AI as a prompt.
[0633] The next generation AI creates a new product and returns it to the server.
[0634] The server repeats this process as necessary.
[0635] Step 7:
[0636] The server then sends the final integrated product to the user's terminal.
[0637] Specific behavior:
[0638] The server converts the aggregated product into the appropriate format (text, images, other media).
[0639] The server sends the result to the user's terminal.
[0640] The user checks the generated data on the terminal and downloads it if necessary.
[0641] Step 8:
[0642] Users use the Generative AI App Store to search, select, purchase, and review new Generative AI.
[0643] Specific behavior:
[0644] Access the generated AI app store from your device.
[0645] Users use the search function to find generative AI.
[0646] The server displays search results, ratings, and reviews.
[0647] Users select the generated AI and make purchases or reviews.
[0648] The server recommends generative AI based on user ratings.
[0649] Through the above steps, the system allows users to efficiently obtain high-quality products.
[0650] Example 1
[0651] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0652] In today's digital service delivery environment, it is difficult for users to centrally manage a wide variety of generative AIs and obtain high-quality results. Furthermore, there is a lack of mechanisms for automatically coordinating work between different generative AIs, which often results in a lack of consistency in the quality of the results. This requires users to manually switch between individual generative AIs, which increases the complexity of the operation. Furthermore, there is an insufficient process for sharing results between generative AIs, making it difficult to improve the quality of the results.
[0653] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0654] In this invention, the server includes: a means for a user to input a prompt and select multiple generation AIs; a means for a terminal to transmit information about the prompt received from the user and the selected generation AI to the server; a means for the server to convert the prompt received from the user into a format suitable for each generation AI and transmit it to the multiple generation AIs; a means for the generation AI to generate a product based on the prompt; a means for the server to receive and integrate the products from each generation AI; and a means for the server to transmit the integrated product to the user's terminal. This allows users to efficiently use multiple generation AIs and obtain high-quality, consistent products. The server also includes a means for managing call and response between the generation AIs and retransmitting results from the first-stage generation AI to the second-stage generation AI. This process further improves the quality of the products. Furthermore, the system includes a means for a user to search for, select, purchase, and review new generation AIs through a generation AI app store, and a means for the server to recommend generation AIs based on ratings and reviews, allowing users to easily find and use the generation AI that best suits their needs.
[0655] A "user" is an entity that uses the system to select a project, input prompts, and select from multiple generation AIs.
[0656] A "prompt" is a text message that a user enters to instruct the generated AI on a specific task.
[0657] "Generative AI" is a type of artificial intelligence that generates specific outcomes based on prompts provided by the user.
[0658] A "terminal" is an electronic device that can be accessed and operated by a user and is used to input and send prompts to a server.
[0659] The "server" is a central processing unit that converts prompts received from users into a format suitable for each generation AI, sends them to the generation AI, and integrates the results to provide them to the user.
[0660] "Products" are the output results that the generative AI creates based on prompts, including text, images, and other digital content.
[0661] "Call and response" is a process of passing results between generation AIs, and is a method of improving the quality of the product by retransmitting the results obtained from the first-stage generation AI to the next-stage generation AI.
[0662] "Generative AI App Store" means an online platform where users can search, select, purchase, and review generative AI.
[0663] "Recommendation" is the act of the server selecting and suggesting an appropriate generation AI based on the user's needs, ratings, and reviews.
[0664] The present invention provides a platform that allows users to efficiently use multiple generation AIs to obtain high-quality products. An embodiment of this system will be described in detail below.
[0665] User Interface
[0666] Users access the platform through a web browser or dedicated application using an internet-connected device (e.g., personal computer, tablet, smartphone, etc.). In this interface, users are presented with a screen to select a specific project and a screen to enter prompts. After selecting a project and entering prompts, users can select the generative AI to use.
[0667] Specific examples
[0668] The user inputs "I want to create the first act of a drama script" and selects "Drama script generation AI." They can also select "Landscape drawing AI" for drawing scenery.
[0669] Prompt and generate AI selection information
[0670] After the user selects a prompt and a generated AI, the device sends this information to the server in a structured data format such as JSON.
[0671] Prompt Translation
[0672] The server converts the prompts received from the device into a format that is easy for the AI to understand. This process is performed using a natural language processing module. The server converts the prompts into a format suitable for each AI and sends them to each AI.
[0673] Specific examples
[0674] The server converts the prompt "Please describe the scene where the main character uses magic for the first time" into "Scene 1: The main character uses magic for the first time in the square. Describe the scene and emotions in detail," and sends this to the drama script generation AI.
[0675] Receiving the generated AI's response
[0676] The generation AI creates a product based on the prompts and sends the results back to the server. The server receives these responses and integrates them. Specifically, it combines the results of the drama script generation AI with the results of the scenery drawing AI.
[0677] Specific examples
[0678] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[0679] Improving product quality through call and response
[0680] The server improves the quality of the generated results by retransmitting the results of the first-stage generation AI to the next-stage generation AI. This process is called call and response, and allows results to be passed between multiple generation AIs.
[0681] Specific examples
[0682] The server sends the results of the drama script generation AI to the scenery drawing AI, which then uses that information to generate new results. For example, the scenery drawing AI might return a result such as "depict the scenery at the moment the main character casts a spell in more detail."
[0683] Integration and delivery of the final product
[0684] The server aggregates the final product and sends it to the user's device, ensuring that the user receives a consistent, high-quality product at once.
[0685] Specific examples
[0686] The server sends the integrated results to the user as a "detailed script for the scene where magic is used in the square" and a "scenery description of that scene."
[0687] Use of the Generative AI App Store
[0688] Users can search, select, purchase, and review new AI generators through the AI generator app store, and the server will recommend the best AI generators to users based on these ratings and reviews.
[0689] Specific examples
[0690] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[0691] As described above, the present invention provides a unified management platform that allows users to efficiently utilize multiple generative AIs and obtain high-quality, consistent results.
[0692] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0693] Step 1:
[0694] The user selects a project and enters a prompt. This information is entered by accessing the platform from the user's device. The user selects the prompt and the generation AI to use.
[0695] Specific actions
[0696] The user accesses the platform's web page using a browser or app. On the project selection screen, they select "Drama Script" and then enter "Please create a scene where the main character uses magic for the first time" on the prompt input screen. They also select "Drama Script Generation AI" and "Landscape Drawing AI" as the generation AIs to use.
[0697] input
[0698] The prompt entered by the user, "Please create a scene where the main character uses magic for the first time," and the selection information of the generated AI.
[0699] output
[0700] The user's input data is temporarily saved in the device's internal storage.
[0701] Step 2:
[0702] The terminal transmits the prompt received from the user and information about the selected generated AI to the server.
[0703] Specific actions
[0704] When the user clicks the "Submit" button, the device converts the prompt and AI selection information into JSON format and sends it to the server.
[0705] input
[0706] User-entered prompts and generated AI selection information.
[0707] output
[0708] The data is sent to the server in JSON format.
[0709] Step 3:
[0710] The server converts the prompts received from the user into a format suitable for each generation AI and sends them to multiple generation AIs.
[0711] Specific actions
[0712] The server uses a natural language processing module to analyze the prompt and convert it into an input format for each generation AI. For example, for the "drama script generation AI," it converts it into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," and for the "scenery drawing AI," it converts it into "Describe the scene at the moment the magical light is emitted in the square."
[0713] input
[0714] The prompt and generated AI selection information received by the server in JSON format.
[0715] output
[0716] The prompt is converted into a format suitable for the generation AI and sent to the corresponding generation AI.
[0717] Step 4:
[0718] The generation AI generates artifacts based on prompts.
[0719] Specific actions
[0720] Each AI generator uses its internal algorithm to create a result based on the prompts it receives. For example, the "drama script generator AI" generates a scenario, and the "landscape drawing AI" generates an illustration.
[0721] input
[0722] The prompt converted into a format suitable for generative AI.
[0723] output
[0724] The artifacts (scripts, illustrations, etc.) are generated and sent back to the server.
[0725] Step 5:
[0726] The server receives and integrates the products from each generation AI.
[0727] Specific actions
[0728] The server receives the response data from each generation AI and combines them into a single integrated product, for example, integrating a scene from a drama script with the corresponding landscape description.
[0729] input
[0730] The products received from each generating AI.
[0731] output
[0732] Integrated artifacts (e.g., "A detailed script for the magic scene in the square" and "A landscape description of the scene").
[0733] Step 6:
[0734] The server manages the call and response between the generation AIs and retransmits the results from the first generation AI to the next generation AI in order to improve the quality of the results.
[0735] Specific actions
[0736] The server retransmits the results of the drama script generation AI to the scenery drawing AI, which then generates new products based on that information.
[0737] input
[0738] Results generated by the first-stage generation AI.
[0739] output
[0740] Improved production from next-stage generation AI.
[0741] Step 7:
[0742] The server aggregates the final product and sends it to the user's device.
[0743] Specific actions
[0744] The server integrates the products obtained from all the generation AIs, formats them as the final product, and provides it to the user.
[0745] input
[0746] A product integrated from multiple generation AIs.
[0747] output
[0748] The final product is sent to the user's terminal.
[0749] Step 8:
[0750] Users search, select, purchase, and review new generative AI through the generative AI app store.
[0751] Specific actions
[0752] Users access the app store, search for and purchase a generative AI, and write a review. The server analyzes these ratings and reviews and recommends the best generative AI for the user.
[0753] input
[0754] User ratings and reviews.
[0755] output
[0756] Recommendation information from the server.
[0757] (Application example 1)
[0758] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0759] In conventional design generation systems using generative AI models, it was difficult to effectively integrate multiple generative AI models and easily create high-quality 3D designs and product introduction videos for virtual stores. Furthermore, collaboration between generative AI models was insufficient, preventing sufficient improvement in the quality of the results. Furthermore, users lacked appropriate information when searching for and selecting new generative AI models, making it difficult to find the optimal model.
[0760] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0761] In this invention, the server includes means for a user to input a prompt sentence and select multiple generative AI models, means for the server to convert the prompt sentence received from the user into a format suitable for each generative AI model and send it to the multiple generative AI models, means for the generative AI models to generate products based on the prompt sentence, means for the server to receive and integrate the products from each generative AI model, and means for the operator of the virtual store to generate high-quality 3D designs and product introduction videos. This allows users to effectively use multiple generative AI models and obtain high-quality products at once, making it possible to generate high-quality content in the virtual store.
[0762] "User device" means an electronic device operated by a user, including a smartphone, smart glasses, a head-mounted display, etc.
[0763] A "generative AI model" refers to an artificial intelligence algorithm that automatically creates a specific product based on a prompt entered by a user.
[0764] A "prompt" is a text expression that succinctly summarizes the input information and instructions that a user provides to a generative AI model.
[0765] "Products" are the deliverables that the generative AI model creates based on the prompt, including 3D designs and product introduction videos.
[0766] A "server" is a central processing unit that runs on a network and manages the exchange of data between the user's device and the generative AI model.
[0767] A "virtual store" is a virtual store that displays and sells products and services online.
[0768] "Call and response between generative AI models" is a technique in which generative AI models work together to repeatedly respond, improving the accuracy and quality of the products.
[0769] A "Generative AI Model App Store" is an online platform where users can search for, select, purchase, and review new generative AI models.
[0770] A system for implementing the present invention includes a user terminal, a server, and multiple generative AI models. The user terminal is an electronic device such as a smartphone, smart glasses, or a head-mounted display, and provides an interface for a user to input a prompt sentence and select a generative AI model.
[0771] Hardware and software used
[0772] Hardware: Smartphones (e.g., iPhone, Samsung Galaxy), smart glasses (e.g., Google Glass), head-mounted displays (e.g., Meta Quest)
[0773] Software: iOS / Android app, web server, generative AI model (e.g., OpenAI's ChatGPT, DALL-E)
[0774] System Operation
[0775] 1. User's device:
[0776] Users access the application from their device, select a specific project, enter a prompt, and can select the generative AI model they want to use.
[0777] example:
[0778] I want to create a 3D model of a new product, and I need high-resolution images as well as a promotional video.
[0779] At this time, the user can select product design AI, video generation AI, or landscape drawing AI.
[0780] 2. Server prompt translation:
[0781] The server converts the prompt received from the user into a format suitable for each generative AI model, and then sends the converted prompt to each generative AI model.
[0782] 3. How generative AI models work:
[0783] The generative AI model creates a specified product based on the received prompt. For example, a product design AI generates a 3D model, a video generation AI generates a product introduction video, and a landscape drawing AI generates a background design.
[0784] 4. Server integration process:
[0785] The server receives the results from each generative AI model and integrates them. The integrated results are then sent to the user's device. At this time, a call-and-response function between the generative AI models is also activated, and the results of multiple generative AI models are sequentially linked to improve the quality of the results.
[0786] 5. User Interface:
[0787] The user can again review the product through the application, provide feedback on the product if necessary, and enter further prompts.
[0788] 6. Generative AI model app store:
[0789] Users can use the generative AI model app store to search for, select, purchase, and review new generative AI models, and the server has the function of recommending generative AI models based on ratings and reviews.
[0790] In this way, users can efficiently utilize a variety of generative AI models to easily generate high-quality 3D designs and product introduction videos for virtual stores.
[0791] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0792] Step 1:
[0793] A user accesses the application from their device, selects a specific project, and enters a prompt. The user selects the generative AI model to use, such as product design AI, video generation AI, or landscape drawing AI. The prompt entered might be something like, "I would like to create a 3D model of a new product. I also need a promotional video along with high-resolution images."
[0794] Step 2:
[0795] The terminal transmits the input prompt sentence and information about the selected generative AI model to the server. At this time, the input data from the terminal includes the prompt sentence and the ID information of the generative AI model.
[0796] Step 3:
[0797] The server analyzes the prompt received from the user and converts it into a format suitable for each generative AI model. For example, a prompt such as "I would like to create a 3D model of a new product" is converted into a format suitable for a product design AI to generate a high-resolution 3D model, and the part "We also need an introductory video" is converted into a format suitable for a video generation AI. The converted prompt is then sent to the generative AI model.
[0798] Step 4:
[0799] The generative AI models (product design AI, video generation AI, and landscape drawing AI) create the specified product based on the received prompt. The product design AI generates a high-resolution 3D model, the video generation AI generates a video introducing the product, and the landscape drawing AI generates a background design. Each generative AI model uses its internal algorithm to calculate data based on the prompt.
[0800] Step 5:
[0801] Each generative AI model completes the product and returns the data to the server. For example, a product design AI sends 3D model data, a video generation AI sends a video file, and a landscape drawing AI sends a background image.
[0802] Step 6:
[0803] The server integrates the results received from each generative AI model. It uses an integration algorithm to optimally combine the data, combining the 3D model data, video files, and background images into a single piece of content. During this integration process, a call-and-response function between the generative AI models is utilized to adjust the results to ensure higher quality.
[0804] Step 7:
[0805] The server transmits the integrated product to the user's terminal, which displays the received product and an interface for providing feedback if necessary.
[0806] Step 8:
[0807] When a user uses the generative AI model app store, the user can search, select, purchase, and review new generative AI models, and the server recommends the most suitable generative AI models based on the user's ratings and reviews.
[0808] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0809] The present invention provides a system that allows users to efficiently use multiple AI generators and further improves the quality of the output by recognizing the user's emotional state. The following describes in detail how the present invention is implemented.
[0810] User Interface
[0811] The process begins when a user accesses the platform from their device and logs in. The user selects a project on the interface screen, fills in prompts, and selects the generative AI to use. This interface works in conjunction with an emotion engine that also recognizes the user's emotions.
[0812] example:
[0813] The user inputs "I want to create the first act of a drama script" into the device and selects the "drama script generation AI." They also select the "landscape drawing AI" to draw the scenery. The emotion engine then analyzes the user's voice and facial expressions and recognizes them as "excited."
[0814] Sending prompts and selecting generation AI
[0815] The device sends the prompts entered by the user, information on the selected generation AI, and emotional data obtained from the emotion engine to the server, which then adjusts the prompts according to the user's emotional state.
[0816] example:
[0817] The prompt sent by the user, "Draw a scene where the main character uses magic for the first time," is received by the server, and the emotion engine detects the "excited" state, so the prompt is adjusted slightly.
[0818] Prompt conversion and generation AI transmission
[0819] The server converts the received prompts and emotion data into a format suitable for each generation AI and sends it to each generation AI. Data obtained from the emotion engine is also applied, so the generated content is adjusted to match the user's emotion.
[0820] example:
[0821] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," adds emotional data, and sends it to the drama script generation AI.
[0822] Response of the generating AI and reception of the product
[0823] The server receives responses from each AI generator. The AI generator creates a product based on each prompt and returns it to the server. The server then integrates these products and provides them to the user in the most optimal form.
[0824] example:
[0825] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[0826] Call and response between generative AI
[0827] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the results. This call-and-response process allows the generation AIs to cooperate with each other. Data from the emotion engine is also continuously used.
[0828] example:
[0829] The results of the drama script generation AI are sent to the scenery drawing AI, which then generates a new product based on that scene.
[0830] Integration and delivery of the final product
[0831] The server integrates the results from multiple generative AIs and provides the final result to the user, allowing users to receive high-quality results at once. The result is also further adjusted based on emotional data.
[0832] example:
[0833] The server sends the integrated results to the user as a detailed script for the scene in which the protagonist uses magic in the square and a description of the scenery of that scene. The description of the scene is adjusted to be even more dramatic depending on the user's level of excitement.
[0834] Use of the Generative AI App Store
[0835] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs based on the user's needs. This process also utilizes data from the emotion engine to recommend Generative AIs that are tailored to the user's current emotional state.
[0836] example:
[0837] Users select and install a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project." Users who are in an "excited" state will be suggested generative AIs that will further increase their excitement.
[0838] In this way, the present invention provides a centralized management platform for users to efficiently utilize multiple generative AIs and further improve the quality of the output using an emotion engine.
[0839] The processing flow will be explained below.
[0840] Step 1:
[0841] The user accesses the platform from their device and logs in. Login authentication is performed and the user's personal information is confirmed.
[0842] Specific behavior:
[0843] The user enters their ID and password on the login screen.
[0844] The device sends the login information to the server.
[0845] The server refers to the user database and performs authentication.
[0846] If the authentication is successful, the user will be taken to the dashboard screen.
[0847] Step 2:
[0848] The user selects a project, enters prompts, and selects each generated AI. The emotion engine then analyzes the user's emotional state.
[0849] Specific behavior:
[0850] On the device interface screen, select the desired project from the project list.
[0851] The user enters a prompt and selects the generated AI to use from a drop-down menu or list.
[0852] An emotion engine built into or connected to the device collects and analyzes emotional data from the user's voice and facial expressions.
[0853] The device sends prompts, the generated AI's selection information, and emotion data to the server.
[0854] Step 3:
[0855] The server receives the sent prompt, the selected generation AI information, and the emotion data from the emotion engine, and converts the prompt into a format suitable for each generation AI based on that information.
[0856] Specific behavior:
[0857] The server parses the prompt and converts it into a format to send to the generating AI.
[0858] Prepare to use emotion data obtained from the emotion engine to adjust the content of prompts and productions.
[0859] Create the necessary requests according to the API endpoint of each generation AI.
[0860] Step 4:
[0861] The server sends the converted prompt to each generation AI, requesting it to generate a product.
[0862] Specific behavior:
[0863] The server sends an HTTP request to the generated AI's API endpoint.
[0864] Each generation AI receives a prompt and creates the specified product.
[0865] Step 5:
[0866] Each generation AI creates a product and returns a response to the server, which receives these products.
[0867] Specific behavior:
[0868] Each generation AI generates a product based on a prompt.
[0869] The server receives the HTTP response from the generation AI and obtains the generated data.
[0870] Step 6:
[0871] The products received by the server are further adjusted and integrated through call and response between each generation AI.
[0872] Specific behavior:
[0873] The server sends the result from the first stage generation AI to the next stage generation AI as a prompt.
[0874] The next generation AI creates a new product and returns it to the server.
[0875] Repeat this process as necessary.
[0876] At each stage, the emotion engine data is used to adjust the output.
[0877] Step 7:
[0878] The final integrated product is sent to the user's device by the server, where it is further adjusted based on the emotion data.
[0879] Specific behavior:
[0880] The server converts the aggregated product into the appropriate format (text, images, other media).
[0881] The server further adjusts the product based on data from the emotion engine and sends it to the user's device.
[0882] Step 8:
[0883] Users can review the results they receive, make corrections and provide feedback as needed, and use the Generative AI App Store to find new Generative AI.
[0884] Specific behavior:
[0885] The user checks the output on the terminal.
[0886] Make corrections as needed and send feedback to the server.
[0887] The device accesses the generative AI app store, and the user searches for, selects, and purchases new generative AI.
[0888] The server recommends a generative AI based on user feedback.
[0889] Example 2
[0890] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0891] Conventional generative AI systems make it difficult for users to efficiently use multiple generative AIs, and have limited means for improving the consistency and quality of the results. Furthermore, they often fail to meet user expectations because they are unable to adjust the results to take into account the user's emotional state. The present invention aims to solve these problems and provide high-quality results that respond to the user's emotional state.
[0892] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for converting prompts and emotion data received from the user into a format suitable for each generation AI and transmitting the converted data to the multiple generation AIs, means for the generation AIs to generate products based on the prompts and emotion data, and means for the server to receive and integrate the products from the generation AIs. This makes it possible to adjust prompts based on the user's emotional state and consistently provide high-quality products.
[0893] A "user" is a person or institution that utilizes the system to enter prompts and select a generating AI.
[0894] "Terminal" means an electronic device used by a user to access the system and provide input information.
[0895] The "server" is a central processing unit that processes information received from users, converts it into a format suitable for the generative AI model, integrates the generated results, and provides them to users.
[0896] "Generative AI" is an artificial intelligence model that creates a specified artifact based on user-entered prompts.
[0897] A "prompt" is an instruction that the user enters to specify the desired product for the generation AI.
[0898] A "product" is an output result that the generation AI generates based on a prompt.
[0899] "Emotional data" refers to data that indicates the user's current emotional state and is used to adjust the product.
[0900] The "emotion engine" is a function that analyzes the user's voice, facial expressions, etc. to generate emotional data.
[0901] "Call and response" is a process in which the results obtained from the first-stage generation AI are retransmitted to the next-stage generation AI to improve the quality of the product.
[0902] The "Generative AI App Store" is an online platform where users can search for, select, purchase, and review new generative AI.
[0903] "Recommendation" means that the server suggests a generative AI that is suitable for the user based on ratings and reviews.
[0904] "Format conversion" is the process by which the server converts the prompt and emotion data received from the user into a format suitable for the generative AI.
[0905] "Integration" refers to the server combining the products obtained from each generation AI into one.
[0906] The system of the present invention is a platform that allows users to efficiently utilize multiple generation AIs and improve the quality of the products based on the user's emotional state. The components and specific functions of this system will be described below.
[0907] User Interface
[0908] Users access the platform from their own device and log in. Next, they select a project and enter a prompt for the AI generator. The user then selects the AI generator they want to use. At this time, an emotion engine is connected to the device, which analyzes and recognizes the user's emotional state in real time.
[0909] Example prompt sentence:
[0910] "I want to create the first act of a drama script."
[0911] Select "Drama script generation AI" and also select "Landscape drawing AI" for landscape drawing.
[0912] The emotion engine analyzes the user's voice and facial expressions to detect an "excited" state.
[0913] Sending prompt information and emotion data
[0914] The device sends the prompt text entered by the user, the selected generation AI information, and the emotion data obtained from the emotion engine to the server, which receives this information and prepares it for processing.
[0915] Prompt adjustment and transformation
[0916] The server analyzes the received prompt and emotion data, adjusts the prompt content as needed, and then converts it into the optimal format for each generative AI model before sending it.
[0917] Examples:
[0918] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," adds emotional data about excitement, and sends it to the drama script generation AI.
[0919] Receiving and integrating generative AI responses
[0920] The server receives responses from each AI generator. Each AI generator creates a product based on the prompt and emotion data and sends it back to the server. The server then integrates these products and provides them to the user in the most optimal form.
[0921] Examples:
[0922] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[0923] Call and response between generative AI
[0924] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the product. In this process, collaboration between the generation AIs is realized, and emotion data is also continuously used.
[0925] Examples:
[0926] The server sends the results of the drama script generation AI to the scenery drawing AI, which then generates a new product based on that scene.
[0927] Providing the final product
[0928] The server integrates the results from multiple generative AIs and provides the final result to the user. The result is further adjusted based on the user's emotional data, allowing the user to receive high-quality results all at once.
[0929] Examples:
[0930] The server sends the integrated results to the user as a "detailed script for the scene in which the protagonist uses magic in the square" and a "scenery description of that scene."
[0931] The depiction of the scene is adjusted to become more dramatic depending on the user's level of excitement.
[0932] Use of the Generative AI App Store
[0933] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs that meet the user's needs. Data from the emotion engine is also utilized to suggest Generative AIs that are suited to the user's current emotional state.
[0934] Examples:
[0935] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[0936] For users who are excited, generative AI will be suggested to further increase their excitement.
[0937] In this way, the present invention provides a unified platform that allows users to efficiently utilize multiple generative AI models and provide the highest quality products based on emotion data.
[0938] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0939] Step 1:
[0940] A user accesses the platform from their device and logs in. At this time, the user enters authentication information (username, password) and is granted access to the system. If the entered authentication information is correct, the server starts a session and displays the user interface (UI). This UI includes a project selection screen and a generation AI selection screen.
[0941] Input: Username, Password
[0942] Output: Session starts, UI is displayed
[0943] Specific behavior:
[0944] The user enters their credentials on the login page and clicks the submit button.
[0945] The server verifies the authentication information and, if correct, displays the dashboard screen.
[0946] Step 2:
[0947] The user selects a project on the interface screen and inputs a prompt for the AI to generate. The user also selects the AI to use. The device is connected to an emotion engine that analyzes and recognizes the user's emotional state in real time.
[0948] Input: Project selection, prompt text, AI generation selection
[0949] Output: Project information, prompt information, emotion data
[0950] Specific behavior:
[0951] The user enters "I want to create the first act of a drama script" and selects "Drama script generation AI."
[0952] The emotion engine scans the user's facial expressions to detect "excitement."
[0953] Step 3:
[0954] The device sends the prompt text entered by the user, information on the selected generation AI, and emotion data obtained from the emotion engine to the server, which receives this information and prepares to adjust the prompt text.
[0955] Input: prompt, generated AI information, emotion data
[0956] Output: Send data to the server
[0957] Specific behavior:
[0958] The terminal sends project information, prompts and "excitement state" data to the server.
[0959] Step 4:
[0960] The server analyzes the received prompt and emotion data, adjusts the prompt as needed, and then converts it into a format optimal for each generative AI model and sends it.
[0961] Input: prompt sentence, emotion data
[0962] Output: Adjusted prompt sentence, data sent to the generation AI
[0963] Specific behavior:
[0964] The server translates the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail."
[0965] Data adjusted to a format suitable for the generative AI model is sent to each generative AI.
[0966] Step 5:
[0967] The generation AI generates a product based on the prompt sentence and emotion data and sends it back to the server. The server receives responses from each generation AI and integrates the products.
[0968] Input: Adjusted prompt sentence, emotion data
[0969] Output: Artifacts (e.g. text, images)
[0970] Specific behavior:
[0971] Generative AI generates text and images based on the prompts it receives.
[0972] The server receives and integrates the products returned by the generation AI.
[0973] example:
[0974] Response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience erupts in amazement and cheers."
[0975] The landscape drawing AI responds, "The square is lush and green, and the magic light is blue."
[0976] Step 6:
[0977] The server retransmits the results of the first-stage AI generation to the second-stage AI generation to further improve the quality of the product. Emotion data is also continuously used. This results in a product of improved quality.
[0978] Input: Generation results from the first stage generation AI, emotion data
[0979] Output: Refined product from next stage generation AI
[0980] Specific behavior:
[0981] The server sends the results of the first stage drama script generation AI to the scenery drawing AI.
[0982] The landscape drawing AI generates new products based on the received data and sends them back.
[0983] Step 7:
[0984] The server integrates the results from multiple generative AIs and provides the final result to the user. Emotional data is also taken into account, allowing the user to receive the results in the most optimal way.
[0985] Input: Products from each generation AI, emotion data
[0986] Output: The integrated final product
[0987] Specific behavior:
[0988] The server integrates the results of each generation AI to create the final scene script and scenery description.
[0989] The scene depiction is adjusted to be even more dramatic depending on the user's level of excitement.
[0990] Step 8:
[0991] Users access the generative AI app store to search, select, purchase, and review new generative AIs. The server then recommends generative AIs that meet the user's needs, utilizing emotional data to make suggestions.
[0992] Input: User searches, selections, purchase information, ratings and reviews, sentiment data
[0993] Output: Recommendations and generative AI
[0994] Specific behavior:
[0995] A user searches for a new generative AI in the app store, checks reviews, and then purchases it.
[0996] The server recommends, "This generative AI is suitable for your project."
[0997] Through these specific steps and operations, the system of the present invention provides a high-performance and user-friendly generative AI platform.
[0998] (Application example 2)
[0999] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1000] Conventional generative AI systems lack the ability to customize their outputs based on the user's emotional state. This can result in outputs that do not match the user's expectations or emotional state, resulting in a poor user experience. Furthermore, there is also the issue of insufficient collaboration between generative AIs, resulting in a decline in the quality of the outputs. Furthermore, it is difficult to easily search for and select new generative AIs, preventing users from effectively utilizing a wide range of generative AIs.
[1001] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1002] In this invention, the server includes means for recognizing the emotional state of the user and adjusting the quality of the product based on the emotional data, an emotion engine for analyzing the emotional data, means for managing call and response between the generation AIs and retransmitting results from the first-stage generation AI to the next-stage generation AI to improve the quality of the product, means for the user to search, select, purchase, and review new generation AIs through a generation AI app store, and means for the server to recommend generation AI models based on the ratings and reviews. This makes it possible to provide high-quality products tailored to the user's emotional state, strengthens cooperation between generation AIs, and makes it easier to use new generation AIs.
[1003] "User" refers to the person using the generative AI system to input prompts.
[1004] A "prompt" refers to a sentence in the form of specific instructions or questions that are input to the generating AI.
[1005] "Generative AI" refers to an artificial intelligence model that creates artifacts based on prompts from the user.
[1006] "Server" refers to the computer system that receives prompts from the user, converts them into a format suitable for the generating AI, and then integrates and provides the resulting results to the user.
[1007] "Product" refers to the output data that the generative AI creates based on the prompts.
[1008] "Emotional state" refers to the user's current emotional or mental state, as recognized by the emotion engine.
[1009] "Emotion Engine" refers to a software or hardware system that analyzes a user's emotional state and provides that data.
[1010] "Call and response" refers to the collaborative process of retransmitting the product obtained from the first-stage generation AI to the next-stage generation AI to improve the quality of the product.
[1011] "Generative AI App Store" refers to an online platform where users can search for, select, purchase, and review new generative AI.
[1012] "Recommendation" refers to the act of suggesting the most suitable generative AI based on user ratings and reviews.
[1013] "User's device" refers to the keywords for smartphones, computers, tablets, etc. that users use to operate the generated AI system.
[1014] The present invention provides a system that provides high-quality products by efficiently utilizing multiple generation AIs while taking into account the emotional state of the user. Detailed embodiments of the present invention will be described below.
[1015] System Overview
[1016] The system of the present invention includes means for a user to input a prompt and select a generation AI, means for a server to convert the prompt into a suitable format and send it to the generation AI, means for the generation AI to generate a product based on the prompt, means for integrating the product and providing it to the user, and means for using an emotion engine to analyze the user's emotional state and adjust the product based thereon.
[1017] Hardware and software used
[1018] Hardware
[1019] User devices: smartphones, computers, tablets, etc.
[1020] Camera module: For capturing images to analyze the user's facial expressions
[1021] software
[1022] Emotion recognition engine: Software that analyzes the user's emotional state (e.g., the Hugging Face transformer library)
[1023] Generative AI models: Artificial intelligence models that create artifacts based on prompts (e.g., GPT-3)
[1024] Review analysis engine: Software that analyzes reviews based on artifacts
[1025] System Operation
[1026] 1. User Interface
[1027] Users access the interface from a device such as a smartphone or computer and log in. They then perform a search, select a project, fill in the prompts, and can also select a generative AI model.
[1028] 2. Acquiring Emotion Data
[1029] The camera module of the user's device is used to capture facial images, which are then analyzed by an emotion recognition engine to understand the user's emotional state, for example, determining whether the user is in an "excited" state.
[1030] 3. Sending prompts and selecting generation AI
[1031] The prompts and emotion data entered by the user are sent to the server, which converts them into a suitable format and sends them to the selected generative AI model.
[1032] 4. Product generation and preparation
[1033] The generative AI creates products based on prompts and adjusts the content based on emotional data. For example, if a user is in an excited state and enters the prompt, "I'm looking for a black leather jacket. Can you recommend any products?", the generative AI will generate product recommendations that correspond to the user's excited state.
[1034] 5. Consolidating and Presenting Results
[1035] The server integrates the results from each AI generator and provides it to the user, who then adjusts the results to match the user's emotional state.
[1036] 6. Use of Generative AI App Store
[1037] Users can search, select, purchase, and review new generative AI models through the generative AI app store, and the server will recommend the most suitable generative AI based on user ratings and reviews.
[1038] Specific examples
[1039] If a user types a prompt like "I'm looking for a black leather jacket. Can you recommend some?", and the user's emotional state is recognized as "excited," the following happens:
[1040] 1. Prompt and emotion data acquisition
[1041] Prompt: "I'm looking for a black leather jacket. Can you recommend one?"
[1042] Emotional state: "Excited"
[1043] 2. Product generation
[1044] The generative AI model generates detailed product descriptions and recommendations that match the user's excitement level, such as, "This black leather jacket is trending this season and has received very high ratings."
[1045] In this way, the present invention can provide high quality products based on the user's emotional state.
[1046] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1047] Step 1:
[1048] A user uses a terminal to access the interface and log in. The user selects a project and fills in the prompts. Then, the user selects multiple generative AI models.
[1049] Input: User login information, project selection, prompt text, and selection of generative AI model
[1050] Output: Information about the input prompt and the selected generative AI model
[1051] Specific behavior: For example, the user enters the prompt, "I'm looking for a black leather jacket. Can you recommend some?" and selects the relevant generative AI model.
[1052] Step 2:
[1053] The device's camera module is used to capture the user's facial expression, and the image data is sent to the emotion recognition engine for analysis.
[1054] Input: User's face image
[1055] Output: Parsed user emotional state data
[1056] Specific operation: A facial image is captured with a camera, sent to an emotion recognition engine, and the user's emotional state (e.g., excitement) is analyzed.
[1057] Step 3:
[1058] The prompt sentence and emotional state data are sent to the server, which receives the prompt sentence and emotional state data and converts the prompt sentence into a format suitable for each generative AI model.
[1059] Input: prompt sentence, emotional state data
[1060] Output: A prompt in a format suitable for a generative AI model
[1061] What happens: The server translates the prompt to something like, "I'm looking for a black leather jacket. Can you recommend one? [Emotion: Excited]."
[1062] Step 4:
[1063] The server sends the converted prompt sentence to each generative AI model, which creates a product based on the prompt sentence and returns it to the server.
[1064] Input: A prompt in a format suitable for a generative AI model
[1065] Output: The product of the generative AI model
[1066] Specific operation: The generative AI model generates product information and returns a result such as "This black leather jacket is a trendy item this season and has received very high ratings" to the server.
[1067] Step 5:
[1068] The server receives the outputs from each generative AI model, adjusts them based on the emotional state data, and integrates them.
[1069] Input: Product from generative AI model, emotional state data
[1070] Output: The integrated product
[1071] What it does: The server integrates the generated product details and recommendations based on them, dynamically adjusting them to fit the sentiment.
[1072] Step 6:
[1073] The server sends the integrated product to the user's terminal, and the user checks the received product.
[1074] Input: The integrated product
[1075] Output: The artifact provided to the user
[1076] Specific operation: The server sends a message to the user's device such as "Recommended product: This black leather jacket is a trendy item this season and has received very high ratings."
[1077] Step 7:
[1078] Users search, select, purchase, and review new generative AI models through the generative AI app store, and the server recommends generative AI models based on ratings and reviews.
[1079] Input: User ratings and reviews
[1080] Output: A generative AI model recommended to the user
[1081] Specific operation: The user searches for new generative AI models in the generative AI app store, purchases them, and reviews them. The server then uses this information to suggest the optimal generative AI model.
[1082] Through the above processing steps, this system realizes the operation of an advanced generative AI model that takes into account the user's emotional state, providing high-quality results.
[1083] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1084] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1085] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1086] [Third embodiment]
[1087] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1088] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1089] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1090] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1091] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1092] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1093] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1094] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1095] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1096] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1097] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1098] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1099] The present invention is a system that provides a platform for users to efficiently use multiple generation AIs to obtain high-quality products. The operation of the system that embodies the present invention will be described below.
[1100] User Interface
[1101] When users access the platform from their device, they are presented with a screen where they can select a specific project and enter prompts. In this interface, users can select each generated AI.
[1102] example:
[1103] The user inputs into the terminal, "I want to create the first act of a drama script," and selects the "drama script generation AI." He also selects the "landscape drawing AI" to draw the scenery.
[1104] Sending prompts and selecting generation AI
[1105] The terminal sends the prompts entered by the user and the information of the selected generated AI to the server, which then assigns each generated AI to the task specified by the user.
[1106] example:
[1107] The user sends a prompt, "Draw a scene where the main character uses magic for the first time," and it arrives at the server.
[1108] Prompt conversion and generation AI transmission
[1109] The server converts the received prompt into a format suitable for each generation AI and sends it to each generation AI, so that each generation AI can accurately understand and process the prompt.
[1110] example:
[1111] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," and sends it to the drama script generation AI.
[1112] Response of the generating AI and reception of the product
[1113] The server receives responses from each AI generator. The AI generator creates a product based on each prompt and returns it to the server. The server then integrates these products and provides them to the user in the most optimal form.
[1114] example:
[1115] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[1116] Call and response between generative AI
[1117] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the results. This call-and-response process enables collaboration between the generation AIs.
[1118] example:
[1119] The results of the drama script generation AI are sent to the scenery drawing AI, which then generates a new product based on that scene.
[1120] Integration and delivery of the final product
[1121] The server integrates the results from multiple generative AIs and provides the final result to the user, allowing the user to receive high-quality results all at once.
[1122] example:
[1123] The server sends the integrated results to the user as a "detailed script for the scene in which the protagonist uses magic in the square" and a "scenery description of that scene."
[1124] Use of the Generative AI App Store
[1125] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs based on the user's needs.
[1126] example:
[1127] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[1128] In this way, the present invention provides a centralized management platform that allows users to efficiently utilize multiple generation AIs and obtain optimal products.
[1129] The processing flow will be explained below.
[1130] Step 1:
[1131] The user accesses the platform from their device and logs in. This authenticates the user.
[1132] Specific behavior:
[1133] The user enters their ID and password on the login screen.
[1134] The device sends the login information to the server.
[1135] The server refers to the user database and performs authentication.
[1136] Step 2:
[1137] The user selects a project, fills in the prompts, and selects each generated AI.
[1138] Specific behavior:
[1139] On the device interface screen, select the desired project from the project list.
[1140] The user enters a prompt and selects the generated AI to use from a drop-down menu or list.
[1141] The device sends the selections and inputs to the server.
[1142] Step 3:
[1143] The server processes the received prompts and generated AI selection information and converts them into a format suitable for each generated AI.
[1144] Specific behavior:
[1145] The server parses the prompt and converts it into a format to send to each generated AI.
[1146] The server creates the necessary requests according to the generated AI's API endpoint.
[1147] Step 4:
[1148] The server sends the converted prompt to each generated AI and waits for a response.
[1149] Specific behavior:
[1150] The server sends an HTTP request to each generated AI.
[1151] Each generation AI receives a prompt and creates a product based on it.
[1152] The server waits for a response from the spawning AI.
[1153] Step 5:
[1154] After the products are returned from each generation AI, the server receives them, integrates them, and processes them.
[1155] Specific behavior:
[1156] The server receives the HTTP response from the generated AI.
[1157] Analyzes the received artifacts and prepares them for integration.
[1158] If necessary, call and response between products will be carried out.
[1159] Step 6:
[1160] When there is a call and response between generation AIs, the server sends the results of the first generation to the next generation AI.
[1161] Specific behavior:
[1162] The server sends the result from the first stage generation AI to the next stage generation AI as a prompt.
[1163] The next generation AI creates a new product and returns it to the server.
[1164] The server repeats this process as necessary.
[1165] Step 7:
[1166] The server then sends the final integrated product to the user's terminal.
[1167] Specific behavior:
[1168] The server converts the aggregated product into the appropriate format (text, images, other media).
[1169] The server sends the result to the user's terminal.
[1170] The user checks the generated data on the terminal and downloads it if necessary.
[1171] Step 8:
[1172] Users use the Generative AI App Store to search, select, purchase, and review new Generative AI.
[1173] Specific behavior:
[1174] Access the generated AI app store from your device.
[1175] Users use the search function to find generative AI.
[1176] The server displays search results, ratings, and reviews.
[1177] Users select the generated AI and make purchases or reviews.
[1178] The server recommends generative AI based on user ratings.
[1179] Through the above steps, the system allows users to efficiently obtain high-quality products.
[1180] Example 1
[1181] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1182] In today's digital service delivery environment, it is difficult for users to centrally manage a wide variety of generative AIs and obtain high-quality results. Furthermore, there is a lack of mechanisms for automatically coordinating work between different generative AIs, which often results in a lack of consistency in the quality of the results. This requires users to manually switch between individual generative AIs, which increases the complexity of the operation. Furthermore, there is an insufficient process for sharing results between generative AIs, making it difficult to improve the quality of the results.
[1183] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1184] In this invention, the server includes: a means for a user to input a prompt and select multiple generation AIs; a means for a terminal to transmit information about the prompt received from the user and the selected generation AI to the server; a means for the server to convert the prompt received from the user into a format suitable for each generation AI and transmit it to the multiple generation AIs; a means for the generation AI to generate a product based on the prompt; a means for the server to receive and integrate the products from each generation AI; and a means for the server to transmit the integrated product to the user's terminal. This allows users to efficiently use multiple generation AIs and obtain high-quality, consistent products. The server also includes a means for managing call and response between the generation AIs and retransmitting results from the first-stage generation AI to the second-stage generation AI. This process further improves the quality of the products. Furthermore, the system includes a means for a user to search for, select, purchase, and review new generation AIs through a generation AI app store, and a means for the server to recommend generation AIs based on ratings and reviews, allowing users to easily find and use the generation AI that best suits their needs.
[1185] A "user" is an entity that uses the system to select a project, input prompts, and select from multiple generation AIs.
[1186] A "prompt" is a text message that a user enters to instruct the generated AI on a specific task.
[1187] "Generative AI" is a type of artificial intelligence that generates specific outcomes based on prompts provided by the user.
[1188] A "terminal" is an electronic device that can be accessed and operated by a user and is used to input and send prompts to a server.
[1189] The "server" is a central processing unit that converts prompts received from users into a format suitable for each generation AI, sends them to the generation AI, and integrates the results to provide them to the user.
[1190] "Products" are the output results that the generative AI creates based on prompts, including text, images, and other digital content.
[1191] "Call and response" is a process of passing results between generation AIs, and is a method of improving the quality of the product by retransmitting the results obtained from the first-stage generation AI to the next-stage generation AI.
[1192] "Generative AI App Store" means an online platform where users can search, select, purchase, and review generative AI.
[1193] "Recommendation" is the act of the server selecting and suggesting an appropriate generation AI based on the user's needs, ratings, and reviews.
[1194] The present invention provides a platform that allows users to efficiently use multiple generation AIs to obtain high-quality products. An embodiment of this system will be described in detail below.
[1195] User Interface
[1196] Users access the platform through a web browser or dedicated application using an internet-connected device (e.g., personal computer, tablet, smartphone, etc.). In this interface, users are presented with a screen to select a specific project and a screen to enter prompts. After selecting a project and entering prompts, users can select the generative AI to use.
[1197] Specific examples
[1198] The user inputs "I want to create the first act of a drama script" and selects "Drama script generation AI." They can also select "Landscape drawing AI" for drawing scenery.
[1199] Prompt and generate AI selection information
[1200] After the user selects a prompt and a generated AI, the device sends this information to the server in a structured data format such as JSON.
[1201] Prompt Translation
[1202] The server converts the prompts received from the device into a format that is easy for the AI to understand. This process is performed using a natural language processing module. The server converts the prompts into a format suitable for each AI and sends them to each AI.
[1203] Specific examples
[1204] The server converts the prompt "Please describe the scene where the main character uses magic for the first time" into "Scene 1: The main character uses magic for the first time in the square. Describe the scene and emotions in detail," and sends this to the drama script generation AI.
[1205] Receiving the generated AI's response
[1206] The generation AI creates a product based on the prompts and sends the results back to the server. The server receives these responses and integrates them. Specifically, it combines the results of the drama script generation AI with the results of the scenery drawing AI.
[1207] Specific examples
[1208] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[1209] Improving product quality through call and response
[1210] The server improves the quality of the generated results by retransmitting the results of the first-stage generation AI to the next-stage generation AI. This process is called call and response, and allows results to be passed between multiple generation AIs.
[1211] Specific examples
[1212] The server sends the results of the drama script generation AI to the scenery drawing AI, which then uses that information to generate new results. For example, the scenery drawing AI might return a result such as "depict the scenery at the moment the main character casts a spell in more detail."
[1213] Integration and delivery of the final product
[1214] The server aggregates the final product and sends it to the user's device, ensuring that the user receives a consistent, high-quality product at once.
[1215] Specific examples
[1216] The server sends the integrated results to the user as a "detailed script for the scene where magic is used in the square" and a "scenery description of that scene."
[1217] Use of the Generative AI App Store
[1218] Users can search, select, purchase, and review new AI generators through the AI generator app store, and the server will recommend the best AI generators to users based on these ratings and reviews.
[1219] Specific examples
[1220] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[1221] As described above, the present invention provides a unified management platform that allows users to efficiently utilize multiple generative AIs and obtain high-quality, consistent results.
[1222] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1223] Step 1:
[1224] The user selects a project and enters a prompt. This information is entered by accessing the platform from the user's device. The user selects the prompt and the generation AI to use.
[1225] Specific actions
[1226] The user accesses the platform's web page using a browser or app. On the project selection screen, they select "Drama Script" and then enter "Please create a scene where the main character uses magic for the first time" on the prompt input screen. They also select "Drama Script Generation AI" and "Landscape Drawing AI" as the generation AIs to use.
[1227] input
[1228] The prompt entered by the user, "Please create a scene where the main character uses magic for the first time," and the selection information of the generated AI.
[1229] output
[1230] The user's input data is temporarily saved in the device's internal storage.
[1231] Step 2:
[1232] The terminal transmits the prompt received from the user and information about the selected generated AI to the server.
[1233] Specific actions
[1234] When the user clicks the "Submit" button, the device converts the prompt and AI selection information into JSON format and sends it to the server.
[1235] input
[1236] User-entered prompts and generated AI selection information.
[1237] output
[1238] The data is sent to the server in JSON format.
[1239] Step 3:
[1240] The server converts the prompts received from the user into a format suitable for each generation AI and sends them to multiple generation AIs.
[1241] Specific actions
[1242] The server uses a natural language processing module to analyze the prompt and convert it into an input format for each generation AI. For example, for the "drama script generation AI," it converts it into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," and for the "scenery drawing AI," it converts it into "Describe the scene at the moment the magical light is emitted in the square."
[1243] input
[1244] The prompt and generated AI selection information received by the server in JSON format.
[1245] output
[1246] The prompt is converted into a format suitable for the generation AI and sent to the corresponding generation AI.
[1247] Step 4:
[1248] The generation AI generates artifacts based on prompts.
[1249] Specific actions
[1250] Each AI generator uses its internal algorithm to create a result based on the prompts it receives. For example, the "drama script generator AI" generates a scenario, and the "landscape drawing AI" generates an illustration.
[1251] input
[1252] The prompt converted into a format suitable for generative AI.
[1253] output
[1254] The artifacts (scripts, illustrations, etc.) are generated and sent back to the server.
[1255] Step 5:
[1256] The server receives and integrates the products from each generation AI.
[1257] Specific actions
[1258] The server receives the response data from each generation AI and combines them into a single integrated product, for example, integrating a scene from a drama script with the corresponding landscape description.
[1259] input
[1260] The products received from each generating AI.
[1261] output
[1262] Integrated artifacts (e.g., "A detailed script for the magic scene in the square" and "A landscape description of the scene").
[1263] Step 6:
[1264] The server manages the call and response between the generation AIs and retransmits the results from the first generation AI to the next generation AI in order to improve the quality of the results.
[1265] Specific actions
[1266] The server retransmits the results of the drama script generation AI to the scenery drawing AI, which then generates new products based on that information.
[1267] input
[1268] Results generated by the first-stage generation AI.
[1269] output
[1270] Improved production from next-stage generation AI.
[1271] Step 7:
[1272] The server aggregates the final product and sends it to the user's device.
[1273] Specific actions
[1274] The server integrates the products obtained from all the generation AIs, formats them as the final product, and provides it to the user.
[1275] input
[1276] A product integrated from multiple generation AIs.
[1277] output
[1278] The final product is sent to the user's terminal.
[1279] Step 8:
[1280] Users search, select, purchase, and review new generative AI through the generative AI app store.
[1281] Specific actions
[1282] Users access the app store, search for and purchase a generative AI, and write a review. The server analyzes these ratings and reviews and recommends the best generative AI for the user.
[1283] input
[1284] User ratings and reviews.
[1285] output
[1286] Recommendation information from the server.
[1287] (Application example 1)
[1288] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1289] In conventional design generation systems using generative AI models, it was difficult to effectively integrate multiple generative AI models and easily create high-quality 3D designs and product introduction videos for virtual stores. Furthermore, collaboration between generative AI models was insufficient, preventing sufficient improvement in the quality of the results. Furthermore, users lacked appropriate information when searching for and selecting new generative AI models, making it difficult to find the optimal model.
[1290] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1291] In this invention, the server includes means for a user to input a prompt sentence and select multiple generative AI models, means for the server to convert the prompt sentence received from the user into a format suitable for each generative AI model and send it to the multiple generative AI models, means for the generative AI models to generate products based on the prompt sentence, means for the server to receive and integrate the products from each generative AI model, and means for the operator of the virtual store to generate high-quality 3D designs and product introduction videos. This allows users to effectively use multiple generative AI models and obtain high-quality products at once, making it possible to generate high-quality content in the virtual store.
[1292] "User device" means an electronic device operated by a user, including a smartphone, smart glasses, a head-mounted display, etc.
[1293] A "generative AI model" refers to an artificial intelligence algorithm that automatically creates a specific product based on a prompt entered by a user.
[1294] A "prompt" is a text expression that succinctly summarizes the input information and instructions that a user provides to a generative AI model.
[1295] "Products" are the deliverables that the generative AI model creates based on the prompt, including 3D designs and product introduction videos.
[1296] A "server" is a central processing unit that runs on a network and manages the exchange of data between the user's device and the generative AI model.
[1297] A "virtual store" is a virtual store that displays and sells products and services online.
[1298] "Call and response between generative AI models" is a technique in which generative AI models work together to repeatedly respond, improving the accuracy and quality of the products.
[1299] A "Generative AI Model App Store" is an online platform where users can search for, select, purchase, and review new generative AI models.
[1300] A system for implementing the present invention includes a user terminal, a server, and multiple generative AI models. The user terminal is an electronic device such as a smartphone, smart glasses, or a head-mounted display, and provides an interface for a user to input a prompt sentence and select a generative AI model.
[1301] Hardware and software used
[1302] Hardware: Smartphones (e.g., iPhone, Samsung Galaxy), smart glasses (e.g., Google Glass), head-mounted displays (e.g., Meta Quest)
[1303] Software: iOS / Android app, web server, generative AI model (e.g., OpenAI's ChatGPT, DALL-E)
[1304] System Operation
[1305] 1. User's device:
[1306] Users access the application from their device, select a specific project, enter a prompt, and can select the generative AI model they want to use.
[1307] example:
[1308] I want to create a 3D model of a new product, and I need high-resolution images as well as a promotional video.
[1309] At this time, the user can select product design AI, video generation AI, or landscape drawing AI.
[1310] 2. Server prompt translation:
[1311] The server converts the prompt received from the user into a format suitable for each generative AI model, and then sends the converted prompt to each generative AI model.
[1312] 3. How generative AI models work:
[1313] The generative AI model creates a specified product based on the received prompt. For example, a product design AI generates a 3D model, a video generation AI generates a product introduction video, and a landscape drawing AI generates a background design.
[1314] 4. Server integration process:
[1315] The server receives the results from each generative AI model and integrates them. The integrated results are then sent to the user's device. At this time, a call-and-response function between the generative AI models is also activated, and the results of multiple generative AI models are sequentially linked to improve the quality of the results.
[1316] 5. User Interface:
[1317] The user can again review the product through the application, provide feedback on the product if necessary, and enter further prompts.
[1318] 6. Generative AI model app store:
[1319] Users can use the generative AI model app store to search for, select, purchase, and review new generative AI models, and the server has the function of recommending generative AI models based on ratings and reviews.
[1320] In this way, users can efficiently utilize a variety of generative AI models to easily generate high-quality 3D designs and product introduction videos for virtual stores.
[1321] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1322] Step 1:
[1323] A user accesses the application from their device, selects a specific project, and enters a prompt. The user selects the generative AI model to use, such as product design AI, video generation AI, or landscape drawing AI. The prompt entered might be something like, "I would like to create a 3D model of a new product. I also need a promotional video along with high-resolution images."
[1324] Step 2:
[1325] The terminal transmits the input prompt sentence and information about the selected generative AI model to the server. At this time, the input data from the terminal includes the prompt sentence and the ID information of the generative AI model.
[1326] Step 3:
[1327] The server analyzes the prompt received from the user and converts it into a format suitable for each generative AI model. For example, a prompt such as "I would like to create a 3D model of a new product" is converted into a format suitable for a product design AI to generate a high-resolution 3D model, and the part "We also need an introductory video" is converted into a format suitable for a video generation AI. The converted prompt is then sent to the generative AI model.
[1328] Step 4:
[1329] The generative AI models (product design AI, video generation AI, and landscape drawing AI) create the specified product based on the received prompt. The product design AI generates a high-resolution 3D model, the video generation AI generates a video introducing the product, and the landscape drawing AI generates a background design. Each generative AI model uses its internal algorithm to calculate data based on the prompt.
[1330] Step 5:
[1331] Each generative AI model completes the product and returns the data to the server. For example, a product design AI sends 3D model data, a video generation AI sends a video file, and a landscape drawing AI sends a background image.
[1332] Step 6:
[1333] The server integrates the results received from each generative AI model. It uses an integration algorithm to optimally combine the data, combining the 3D model data, video files, and background images into a single piece of content. During this integration process, a call-and-response function between the generative AI models is utilized to adjust the results to ensure higher quality.
[1334] Step 7:
[1335] The server transmits the integrated product to the user's terminal, which displays the received product and an interface for providing feedback if necessary.
[1336] Step 8:
[1337] When a user uses the generative AI model app store, the user can search, select, purchase, and review new generative AI models, and the server recommends the most suitable generative AI models based on the user's ratings and reviews.
[1338] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1339] The present invention provides a system that allows users to efficiently use multiple AI generators and further improves the quality of the output by recognizing the user's emotional state. The following describes in detail how the present invention is implemented.
[1340] User Interface
[1341] The process begins when a user accesses the platform from their device and logs in. The user selects a project on the interface screen, fills in prompts, and selects the generative AI to use. This interface works in conjunction with an emotion engine that also recognizes the user's emotions.
[1342] example:
[1343] The user inputs "I want to create the first act of a drama script" into the device and selects the "drama script generation AI." They also select the "landscape drawing AI" to draw the scenery. The emotion engine then analyzes the user's voice and facial expressions and recognizes them as "excited."
[1344] Sending prompts and selecting generation AI
[1345] The device sends the prompts entered by the user, information on the selected generation AI, and emotional data obtained from the emotion engine to the server, which then adjusts the prompts according to the user's emotional state.
[1346] example:
[1347] The prompt sent by the user, "Draw a scene where the main character uses magic for the first time," is received by the server, and the emotion engine detects the "excited" state, so the prompt is adjusted slightly.
[1348] Prompt conversion and generation AI transmission
[1349] The server converts the received prompts and emotion data into a format suitable for each generation AI and sends it to each generation AI. Data obtained from the emotion engine is also applied, so the generated content is adjusted to match the user's emotion.
[1350] example:
[1351] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," adds emotional data, and sends it to the drama script generation AI.
[1352] Response of the generating AI and reception of the product
[1353] The server receives responses from each AI generator. The AI generator creates a product based on each prompt and returns it to the server. The server then integrates these products and provides them to the user in the most optimal form.
[1354] example:
[1355] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[1356] Call and response between generative AI
[1357] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the results. This call-and-response process allows the generation AIs to cooperate with each other. Data from the emotion engine is also continuously used.
[1358] example:
[1359] The results of the drama script generation AI are sent to the scenery drawing AI, which then generates a new product based on that scene.
[1360] Integration and delivery of the final product
[1361] The server integrates the results from multiple generative AIs and provides the final result to the user, allowing users to receive high-quality results at once. The result is also further adjusted based on emotional data.
[1362] example:
[1363] The server sends the integrated results to the user as a detailed script for the scene in which the protagonist uses magic in the square and a description of the scenery of that scene. The description of the scene is adjusted to be even more dramatic depending on the user's level of excitement.
[1364] Use of the Generative AI App Store
[1365] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs based on the user's needs. This process also utilizes data from the emotion engine to recommend Generative AIs that are tailored to the user's current emotional state.
[1366] example:
[1367] Users select and install a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project." Users who are in an "excited" state will be suggested generative AIs that will further increase their excitement.
[1368] In this way, the present invention provides a centralized management platform for users to efficiently utilize multiple generative AIs and further improve the quality of the output using an emotion engine.
[1369] The processing flow will be explained below.
[1370] Step 1:
[1371] The user accesses the platform from their device and logs in. Login authentication is performed and the user's personal information is confirmed.
[1372] Specific behavior:
[1373] The user enters their ID and password on the login screen.
[1374] The device sends the login information to the server.
[1375] The server refers to the user database and performs authentication.
[1376] If the authentication is successful, the user will be taken to the dashboard screen.
[1377] Step 2:
[1378] The user selects a project, enters prompts, and selects each generated AI. The emotion engine then analyzes the user's emotional state.
[1379] Specific behavior:
[1380] On the device interface screen, select the desired project from the project list.
[1381] The user enters a prompt and selects the generated AI to use from a drop-down menu or list.
[1382] An emotion engine built into or connected to the device collects and analyzes emotional data from the user's voice and facial expressions.
[1383] The device sends prompts, the generated AI's selection information, and emotion data to the server.
[1384] Step 3:
[1385] The server receives the sent prompt, the selected generation AI information, and the emotion data from the emotion engine, and converts the prompt into a format suitable for each generation AI based on that information.
[1386] Specific behavior:
[1387] The server parses the prompt and converts it into a format to send to the generating AI.
[1388] Prepare to use emotion data obtained from the emotion engine to adjust the content of prompts and productions.
[1389] Create the necessary requests according to the API endpoint of each generation AI.
[1390] Step 4:
[1391] The server sends the converted prompt to each generation AI, requesting it to generate a product.
[1392] Specific behavior:
[1393] The server sends an HTTP request to the generated AI's API endpoint.
[1394] Each generation AI receives a prompt and creates the specified product.
[1395] Step 5:
[1396] Each generation AI creates a product and returns a response to the server, which receives these products.
[1397] Specific behavior:
[1398] Each generation AI generates a product based on a prompt.
[1399] The server receives the HTTP response from the generation AI and obtains the generated data.
[1400] Step 6:
[1401] The products received by the server are further adjusted and integrated through call and response between each generation AI.
[1402] Specific behavior:
[1403] The server sends the result from the first stage generation AI to the next stage generation AI as a prompt.
[1404] The next generation AI creates a new product and returns it to the server.
[1405] Repeat this process as necessary.
[1406] At each stage, the emotion engine data is used to adjust the output.
[1407] Step 7:
[1408] The final integrated product is sent to the user's device by the server, where it is further adjusted based on the emotion data.
[1409] Specific behavior:
[1410] The server converts the aggregated product into the appropriate format (text, images, other media).
[1411] The server further adjusts the product based on data from the emotion engine and sends it to the user's device.
[1412] Step 8:
[1413] Users can review the results they receive, make corrections and provide feedback as needed, and use the Generative AI App Store to find new Generative AI.
[1414] Specific behavior:
[1415] The user checks the output on the terminal.
[1416] Make corrections as needed and send feedback to the server.
[1417] The device accesses the generative AI app store, and the user searches for, selects, and purchases new generative AI.
[1418] The server recommends a generative AI based on user feedback.
[1419] Example 2
[1420] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1421] Conventional generative AI systems make it difficult for users to efficiently use multiple generative AIs, and have limited means for improving the consistency and quality of the results. Furthermore, they often fail to meet user expectations because they are unable to adjust the results to take into account the user's emotional state. The present invention aims to solve these problems and provide high-quality results that respond to the user's emotional state.
[1422] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for converting prompts and emotion data received from the user into a format suitable for each generation AI and transmitting the converted data to the multiple generation AIs, means for the generation AIs to generate products based on the prompts and emotion data, and means for the server to receive and integrate the products from the generation AIs. This makes it possible to adjust prompts based on the user's emotional state and consistently provide high-quality products.
[1423] A "user" is a person or institution that utilizes the system to enter prompts and select a generating AI.
[1424] "Terminal" means an electronic device used by a user to access the system and provide input information.
[1425] The "server" is a central processing unit that processes information received from users, converts it into a format suitable for the generative AI model, integrates the generated results, and provides them to users.
[1426] "Generative AI" is an artificial intelligence model that creates a specified artifact based on user-entered prompts.
[1427] A "prompt" is an instruction that the user enters to specify the desired product for the generation AI.
[1428] A "product" is an output result that the generation AI generates based on a prompt.
[1429] "Emotional data" refers to data that indicates the user's current emotional state and is used to adjust the product.
[1430] The "emotion engine" is a function that analyzes the user's voice, facial expressions, etc. to generate emotional data.
[1431] "Call and response" is a process in which the results obtained from the first-stage generation AI are retransmitted to the next-stage generation AI to improve the quality of the product.
[1432] The "Generative AI App Store" is an online platform where users can search for, select, purchase, and review new generative AI.
[1433] "Recommendation" means that the server suggests a generative AI that is suitable for the user based on ratings and reviews.
[1434] "Format conversion" is the process by which the server converts the prompt and emotion data received from the user into a format suitable for the generative AI.
[1435] "Integration" refers to the server combining the products obtained from each generation AI into one.
[1436] The system of the present invention is a platform that allows users to efficiently utilize multiple generation AIs and improve the quality of the products based on the user's emotional state. The components and specific functions of this system will be described below.
[1437] User Interface
[1438] Users access the platform from their own device and log in. Next, they select a project and enter a prompt for the AI generator. The user then selects the AI generator they want to use. At this time, an emotion engine is connected to the device, which analyzes and recognizes the user's emotional state in real time.
[1439] Example prompt sentence:
[1440] "I want to create the first act of a drama script."
[1441] Select "Drama script generation AI" and also select "Landscape drawing AI" for landscape drawing.
[1442] The emotion engine analyzes the user's voice and facial expressions to detect an "excited" state.
[1443] Sending prompt information and emotion data
[1444] The device sends the prompt text entered by the user, the selected generation AI information, and the emotion data obtained from the emotion engine to the server, which receives this information and prepares it for processing.
[1445] Prompt adjustment and transformation
[1446] The server analyzes the received prompt and emotion data, adjusts the prompt content as needed, and then converts it into the optimal format for each generative AI model before sending it.
[1447] Examples:
[1448] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," adds emotional data about excitement, and sends it to the drama script generation AI.
[1449] Receiving and integrating generative AI responses
[1450] The server receives responses from each AI generator. Each AI generator creates a product based on the prompt and emotion data and sends it back to the server. The server then integrates these products and provides them to the user in the most optimal form.
[1451] Examples:
[1452] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[1453] Call and response between generative AI
[1454] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the product. In this process, collaboration between the generation AIs is realized, and emotion data is also continuously used.
[1455] Examples:
[1456] The server sends the results of the drama script generation AI to the scenery drawing AI, which then generates a new product based on that scene.
[1457] Providing the final product
[1458] The server integrates the results from multiple generative AIs and provides the final result to the user. The result is further adjusted based on the user's emotional data, allowing the user to receive high-quality results all at once.
[1459] Examples:
[1460] The server sends the integrated results to the user as a "detailed script for the scene in which the protagonist uses magic in the square" and a "scenery description of that scene."
[1461] The depiction of the scene is adjusted to become more dramatic depending on the user's level of excitement.
[1462] Use of the Generative AI App Store
[1463] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs that meet the user's needs. Data from the emotion engine is also utilized to suggest Generative AIs that are suited to the user's current emotional state.
[1464] Examples:
[1465] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[1466] For users who are excited, generative AI will be suggested to further increase their excitement.
[1467] In this way, the present invention provides a unified platform that allows users to efficiently utilize multiple generative AI models and provide the highest quality products based on emotion data.
[1468] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1469] Step 1:
[1470] A user accesses the platform from their device and logs in. At this time, the user enters authentication information (username, password) and is granted access to the system. If the entered authentication information is correct, the server starts a session and displays the user interface (UI). This UI includes a project selection screen and a generation AI selection screen.
[1471] Input: Username, Password
[1472] Output: Session starts, UI is displayed
[1473] Specific behavior:
[1474] The user enters their credentials on the login page and clicks the submit button.
[1475] The server verifies the authentication information and, if correct, displays the dashboard screen.
[1476] Step 2:
[1477] The user selects a project on the interface screen and inputs a prompt for the AI to generate. The user also selects the AI to use. The device is connected to an emotion engine that analyzes and recognizes the user's emotional state in real time.
[1478] Input: Project selection, prompt text, AI generation selection
[1479] Output: Project information, prompt information, emotion data
[1480] Specific behavior:
[1481] The user enters "I want to create the first act of a drama script" and selects "Drama script generation AI."
[1482] The emotion engine scans the user's facial expressions to detect "excitement."
[1483] Step 3:
[1484] The device sends the prompt text entered by the user, information on the selected generation AI, and emotion data obtained from the emotion engine to the server, which receives this information and prepares to adjust the prompt text.
[1485] Input: prompt, generated AI information, emotion data
[1486] Output: Send data to the server
[1487] Specific behavior:
[1488] The terminal sends project information, prompts and "excitement state" data to the server.
[1489] Step 4:
[1490] The server analyzes the received prompt and emotion data, adjusts the prompt as needed, and then converts it into a format optimal for each generative AI model and sends it.
[1491] Input: prompt sentence, emotion data
[1492] Output: Adjusted prompt sentence, data sent to the generation AI
[1493] Specific behavior:
[1494] The server translates the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail."
[1495] Data adjusted to a format suitable for the generative AI model is sent to each generative AI.
[1496] Step 5:
[1497] The generation AI generates a product based on the prompt sentence and emotion data and sends it back to the server. The server receives responses from each generation AI and integrates the products.
[1498] Input: Adjusted prompt sentence, emotion data
[1499] Output: Artifacts (e.g. text, images)
[1500] Specific behavior:
[1501] Generative AI generates text and images based on the prompts it receives.
[1502] The server receives and integrates the products returned by the generation AI.
[1503] example:
[1504] Response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience erupts in amazement and cheers."
[1505] The landscape drawing AI responds, "The square is lush and green, and the magic light is blue."
[1506] Step 6:
[1507] The server retransmits the results of the first-stage AI generation to the second-stage AI generation to further improve the quality of the product. Emotion data is also continuously used. This results in a product of improved quality.
[1508] Input: Generation results from the first stage generation AI, emotion data
[1509] Output: Refined product from next stage generation AI
[1510] Specific behavior:
[1511] The server sends the results of the first stage drama script generation AI to the scenery drawing AI.
[1512] The landscape drawing AI generates new products based on the received data and sends them back.
[1513] Step 7:
[1514] The server integrates the results from multiple generative AIs and provides the final result to the user. Emotional data is also taken into account, allowing the user to receive the results in the most optimal way.
[1515] Input: Products from each generation AI, emotion data
[1516] Output: The integrated final product
[1517] Specific behavior:
[1518] The server integrates the results of each generation AI to create the final scene script and scenery description.
[1519] The scene depiction is adjusted to be even more dramatic depending on the user's level of excitement.
[1520] Step 8:
[1521] Users access the generative AI app store to search, select, purchase, and review new generative AIs. The server then recommends generative AIs that meet the user's needs, utilizing emotional data to make suggestions.
[1522] Input: User searches, selections, purchase information, ratings and reviews, sentiment data
[1523] Output: Recommendations and generative AI
[1524] Specific behavior:
[1525] A user searches for a new generative AI in the app store, checks reviews, and then purchases it.
[1526] The server recommends, "This generative AI is suitable for your project."
[1527] Through these specific steps and operations, the system of the present invention provides a high-performance and user-friendly generative AI platform.
[1528] (Application example 2)
[1529] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1530] Conventional generative AI systems lack the ability to customize their outputs based on the user's emotional state. This can result in outputs that do not match the user's expectations or emotional state, resulting in a poor user experience. Furthermore, there is also the issue of insufficient collaboration between generative AIs, resulting in a decline in the quality of the outputs. Furthermore, it is difficult to easily search for and select new generative AIs, preventing users from effectively utilizing a wide range of generative AIs.
[1531] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1532] In this invention, the server includes means for recognizing the emotional state of the user and adjusting the quality of the product based on the emotional data, an emotion engine for analyzing the emotional data, means for managing call and response between the generation AIs and retransmitting results from the first-stage generation AI to the next-stage generation AI to improve the quality of the product, means for the user to search, select, purchase, and review new generation AIs through a generation AI app store, and means for the server to recommend generation AI models based on the ratings and reviews. This makes it possible to provide high-quality products tailored to the user's emotional state, strengthens cooperation between generation AIs, and makes it easier to use new generation AIs.
[1533] "User" refers to the person using the generative AI system to input prompts.
[1534] A "prompt" refers to a sentence in the form of specific instructions or questions that are input to the generating AI.
[1535] "Generative AI" refers to an artificial intelligence model that creates artifacts based on prompts from the user.
[1536] "Server" refers to the computer system that receives prompts from the user, converts them into a format suitable for the generating AI, and then integrates and provides the resulting results to the user.
[1537] "Product" refers to the output data that the generative AI creates based on the prompts.
[1538] "Emotional state" refers to the user's current emotional or mental state, as recognized by the emotion engine.
[1539] "Emotion Engine" refers to a software or hardware system that analyzes a user's emotional state and provides that data.
[1540] "Call and response" refers to the collaborative process of retransmitting the product obtained from the first-stage generation AI to the next-stage generation AI to improve the quality of the product.
[1541] "Generative AI App Store" refers to an online platform where users can search for, select, purchase, and review new generative AI.
[1542] "Recommendation" refers to the act of suggesting the most suitable generative AI based on user ratings and reviews.
[1543] "User's device" refers to the keywords for smartphones, computers, tablets, etc. that users use to operate the generated AI system.
[1544] The present invention provides a system that provides high-quality products by efficiently utilizing multiple generation AIs while taking into account the emotional state of the user. Detailed embodiments of the present invention will be described below.
[1545] System Overview
[1546] The system of the present invention includes means for a user to input a prompt and select a generation AI, means for a server to convert the prompt into a suitable format and send it to the generation AI, means for the generation AI to generate a product based on the prompt, means for integrating the product and providing it to the user, and means for using an emotion engine to analyze the user's emotional state and adjust the product based thereon.
[1547] Hardware and software used
[1548] Hardware
[1549] User devices: smartphones, computers, tablets, etc.
[1550] Camera module: For capturing images to analyze the user's facial expressions
[1551] software
[1552] Emotion recognition engine: Software that analyzes the user's emotional state (e.g., the Hugging Face transformer library)
[1553] Generative AI models: Artificial intelligence models that create artifacts based on prompts (e.g., GPT-3)
[1554] Review analysis engine: Software that analyzes reviews based on artifacts
[1555] System Operation
[1556] 1. User Interface
[1557] Users access the interface from a device such as a smartphone or computer and log in. They then perform a search, select a project, fill in the prompts, and can also select a generative AI model.
[1558] 2. Acquiring Emotion Data
[1559] The camera module of the user's device is used to capture facial images, which are then analyzed by an emotion recognition engine to understand the user's emotional state, for example, determining whether the user is in an "excited" state.
[1560] 3. Sending prompts and selecting generation AI
[1561] The prompts and emotion data entered by the user are sent to the server, which converts them into a suitable format and sends them to the selected generative AI model.
[1562] 4. Product generation and preparation
[1563] The generative AI creates products based on prompts and adjusts the content based on emotional data. For example, if a user is in an excited state and enters the prompt, "I'm looking for a black leather jacket. Can you recommend any products?", the generative AI will generate product recommendations that correspond to the user's excited state.
[1564] 5. Consolidating and Presenting Results
[1565] The server integrates the results from each AI generator and provides it to the user, who then adjusts the results to match the user's emotional state.
[1566] 6. Use of Generative AI App Store
[1567] Users can search, select, purchase, and review new generative AI models through the generative AI app store, and the server will recommend the most suitable generative AI based on user ratings and reviews.
[1568] Specific examples
[1569] If a user types a prompt like "I'm looking for a black leather jacket. Can you recommend some?", and the user's emotional state is recognized as "excited," the following happens:
[1570] 1. Prompt and emotion data acquisition
[1571] Prompt: "I'm looking for a black leather jacket. Can you recommend one?"
[1572] Emotional state: "Excited"
[1573] 2. Product generation
[1574] The generative AI model generates detailed product descriptions and recommendations that match the user's excitement level, such as, "This black leather jacket is trending this season and has received very high ratings."
[1575] In this way, the present invention can provide high quality products based on the user's emotional state.
[1576] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1577] Step 1:
[1578] A user uses a terminal to access the interface and log in. The user selects a project and fills in the prompts. Then, the user selects multiple generative AI models.
[1579] Input: User login information, project selection, prompt text, and selection of generative AI model
[1580] Output: Information about the input prompt and the selected generative AI model
[1581] Specific behavior: For example, the user enters the prompt, "I'm looking for a black leather jacket. Can you recommend some?" and selects the relevant generative AI model.
[1582] Step 2:
[1583] The device's camera module is used to capture the user's facial expression, and the image data is sent to the emotion recognition engine for analysis.
[1584] Input: User's face image
[1585] Output: Parsed user emotional state data
[1586] Specific operation: A facial image is captured with a camera, sent to an emotion recognition engine, and the user's emotional state (e.g., excitement) is analyzed.
[1587] Step 3:
[1588] The prompt sentence and emotional state data are sent to the server, which receives the prompt sentence and emotional state data and converts the prompt sentence into a format suitable for each generative AI model.
[1589] Input: prompt sentence, emotional state data
[1590] Output: A prompt in a format suitable for a generative AI model
[1591] What happens: The server translates the prompt to something like, "I'm looking for a black leather jacket. Can you recommend one? [Emotion: Excited]."
[1592] Step 4:
[1593] The server sends the converted prompt sentence to each generative AI model, which creates a product based on the prompt sentence and returns it to the server.
[1594] Input: A prompt in a format suitable for a generative AI model
[1595] Output: The product of the generative AI model
[1596] Specific operation: The generative AI model generates product information and returns a result such as "This black leather jacket is a trendy item this season and has received very high ratings" to the server.
[1597] Step 5:
[1598] The server receives the outputs from each generative AI model, adjusts them based on the emotional state data, and integrates them.
[1599] Input: Product from generative AI model, emotional state data
[1600] Output: The integrated product
[1601] What it does: The server integrates the generated product details and recommendations based on them, dynamically adjusting them to fit the sentiment.
[1602] Step 6:
[1603] The server sends the integrated product to the user's terminal, and the user checks the received product.
[1604] Input: The integrated product
[1605] Output: The artifact provided to the user
[1606] Specific operation: The server sends a message to the user's device such as "Recommended product: This black leather jacket is a trendy item this season and has received very high ratings."
[1607] Step 7:
[1608] Users search, select, purchase, and review new generative AI models through the generative AI app store, and the server recommends generative AI models based on ratings and reviews.
[1609] Input: User ratings and reviews
[1610] Output: A generative AI model recommended to the user
[1611] Specific operation: The user searches for new generative AI models in the generative AI app store, purchases them, and reviews them. The server then uses this information to suggest the optimal generative AI model.
[1612] Through the above processing steps, this system realizes the operation of an advanced generative AI model that takes into account the user's emotional state, providing high-quality results.
[1613] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1614] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1615] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1616] [Fourth embodiment]
[1617] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1618] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1619] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1620] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1621] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1622] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1623] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1624] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1625] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1626] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1627] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1628] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1629] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1630] The present invention is a system that provides a platform for users to efficiently use multiple generation AIs to obtain high-quality products. The operation of the system that embodies the present invention will be described below.
[1631] User Interface
[1632] When users access the platform from their device, they are presented with a screen where they can select a specific project and enter prompts. In this interface, users can select each generated AI.
[1633] example:
[1634] The user inputs into the terminal, "I want to create the first act of a drama script," and selects the "drama script generation AI." He also selects the "landscape drawing AI" to draw the scenery.
[1635] Sending prompts and selecting generation AI
[1636] The terminal sends the prompts entered by the user and the information of the selected generated AI to the server, which then assigns each generated AI to the task specified by the user.
[1637] example:
[1638] The user sends a prompt, "Draw a scene where the main character uses magic for the first time," and it arrives at the server.
[1639] Prompt conversion and generation AI transmission
[1640] The server converts the received prompt into a format suitable for each generation AI and sends it to each generation AI, so that each generation AI can accurately understand and process the prompt.
[1641] example:
[1642] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," and sends it to the drama script generation AI.
[1643] Response of the generating AI and reception of the product
[1644] The server receives responses from each AI generator. The AI generator creates a product based on each prompt and returns it to the server. The server then integrates these products and provides them to the user in the most optimal form.
[1645] example:
[1646] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[1647] Call and response between generative AI
[1648] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the results. This call-and-response process enables collaboration between the generation AIs.
[1649] example:
[1650] The results of the drama script generation AI are sent to the scenery drawing AI, which then generates a new product based on that scene.
[1651] Integration and delivery of the final product
[1652] The server integrates the results from multiple generative AIs and provides the final result to the user, allowing the user to receive high-quality results all at once.
[1653] example:
[1654] The server sends the integrated results to the user as a "detailed script for the scene in which the protagonist uses magic in the square" and a "scenery description of that scene."
[1655] Use of the Generative AI App Store
[1656] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs based on the user's needs.
[1657] example:
[1658] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[1659] In this way, the present invention provides a centralized management platform that allows users to efficiently utilize multiple generation AIs and obtain optimal products.
[1660] The processing flow will be explained below.
[1661] Step 1:
[1662] The user accesses the platform from their device and logs in. This authenticates the user.
[1663] Specific behavior:
[1664] The user enters their ID and password on the login screen.
[1665] The device sends the login information to the server.
[1666] The server refers to the user database and performs authentication.
[1667] Step 2:
[1668] The user selects a project, fills in the prompts, and selects each generated AI.
[1669] Specific behavior:
[1670] On the device interface screen, select the desired project from the project list.
[1671] The user enters a prompt and selects the generated AI to use from a drop-down menu or list.
[1672] The device sends the selections and inputs to the server.
[1673] Step 3:
[1674] The server processes the received prompts and generated AI selection information and converts them into a format suitable for each generated AI.
[1675] Specific behavior:
[1676] The server parses the prompt and converts it into a format to send to each generated AI.
[1677] The server creates the necessary requests according to the generated AI's API endpoint.
[1678] Step 4:
[1679] The server sends the converted prompt to each generated AI and waits for a response.
[1680] Specific behavior:
[1681] The server sends an HTTP request to each generated AI.
[1682] Each generation AI receives a prompt and creates a product based on it.
[1683] The server waits for a response from the spawning AI.
[1684] Step 5:
[1685] After the products are returned from each generation AI, the server receives them, integrates them, and processes them.
[1686] Specific behavior:
[1687] The server receives the HTTP response from the generated AI.
[1688] Analyzes the received artifacts and prepares them for integration.
[1689] If necessary, call and response between products will be carried out.
[1690] Step 6:
[1691] When there is a call and response between generation AIs, the server sends the results of the first generation to the next generation AI.
[1692] Specific behavior:
[1693] The server sends the result from the first stage generation AI to the next stage generation AI as a prompt.
[1694] The next generation AI creates a new product and returns it to the server.
[1695] The server repeats this process as necessary.
[1696] Step 7:
[1697] The server then sends the final integrated product to the user's terminal.
[1698] Specific behavior:
[1699] The server converts the aggregated product into the appropriate format (text, images, other media).
[1700] The server sends the result to the user's terminal.
[1701] The user checks the generated data on the terminal and downloads it if necessary.
[1702] Step 8:
[1703] Users use the Generative AI App Store to search, select, purchase, and review new Generative AI.
[1704] Specific behavior:
[1705] Access the generated AI app store from your device.
[1706] Users use the search function to find generative AI.
[1707] The server displays search results, ratings, and reviews.
[1708] Users select the generated AI and make purchases or reviews.
[1709] The server recommends generative AI based on user ratings.
[1710] Through the above steps, the system allows users to efficiently obtain high-quality products.
[1711] Example 1
[1712] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1713] In today's digital service delivery environment, it is difficult for users to centrally manage a wide variety of generative AIs and obtain high-quality results. Furthermore, there is a lack of mechanisms for automatically coordinating work between different generative AIs, which often results in a lack of consistency in the quality of the results. This requires users to manually switch between individual generative AIs, which increases the complexity of the operation. Furthermore, there is an insufficient process for sharing results between generative AIs, making it difficult to improve the quality of the results.
[1714] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1715] In this invention, the server includes: a means for a user to input a prompt and select multiple generation AIs; a means for a terminal to transmit information about the prompt received from the user and the selected generation AI to the server; a means for the server to convert the prompt received from the user into a format suitable for each generation AI and transmit it to the multiple generation AIs; a means for the generation AI to generate a product based on the prompt; a means for the server to receive and integrate the products from each generation AI; and a means for the server to transmit the integrated product to the user's terminal. This allows users to efficiently use multiple generation AIs and obtain high-quality, consistent products. The server also includes a means for managing call and response between the generation AIs and retransmitting results from the first-stage generation AI to the second-stage generation AI. This process further improves the quality of the products. Furthermore, the system includes a means for a user to search for, select, purchase, and review new generation AIs through a generation AI app store, and a means for the server to recommend generation AIs based on ratings and reviews, allowing users to easily find and use the generation AI that best suits their needs.
[1716] A "user" is an entity that uses the system to select a project, input prompts, and select from multiple generation AIs.
[1717] A "prompt" is a text message that a user enters to instruct the generated AI on a specific task.
[1718] "Generative AI" is a type of artificial intelligence that generates specific outcomes based on prompts provided by the user.
[1719] A "terminal" is an electronic device that can be accessed and operated by a user and is used to input and send prompts to a server.
[1720] The "server" is a central processing unit that converts prompts received from users into a format suitable for each generation AI, sends them to the generation AI, and integrates the results to provide them to the user.
[1721] "Products" are the output results that the generative AI creates based on prompts, including text, images, and other digital content.
[1722] "Call and response" is a process of passing results between generation AIs, and is a method of improving the quality of the product by retransmitting the results obtained from the first-stage generation AI to the next-stage generation AI.
[1723] "Generative AI App Store" means an online platform where users can search, select, purchase, and review generative AI.
[1724] "Recommendation" is the act of the server selecting and suggesting an appropriate generation AI based on the user's needs, ratings, and reviews.
[1725] The present invention provides a platform that allows users to efficiently use multiple generation AIs to obtain high-quality products. An embodiment of this system will be described in detail below.
[1726] User Interface
[1727] Users access the platform through a web browser or dedicated application using an internet-connected device (e.g., personal computer, tablet, smartphone, etc.). In this interface, users are presented with a screen to select a specific project and a screen to enter prompts. After selecting a project and entering prompts, users can select the generative AI to use.
[1728] Specific examples
[1729] The user inputs "I want to create the first act of a drama script" and selects "Drama script generation AI." They can also select "Landscape drawing AI" for drawing scenery.
[1730] Prompt and generate AI selection information
[1731] After the user selects a prompt and a generated AI, the device sends this information to the server in a structured data format such as JSON.
[1732] Prompt Translation
[1733] The server converts the prompts received from the device into a format that is easy for the AI to understand. This process is performed using a natural language processing module. The server converts the prompts into a format suitable for each AI and sends them to each AI.
[1734] Specific examples
[1735] The server converts the prompt "Please describe the scene where the main character uses magic for the first time" into "Scene 1: The main character uses magic for the first time in the square. Describe the scene and emotions in detail," and sends this to the drama script generation AI.
[1736] Receiving the generated AI's response
[1737] The generation AI creates a product based on the prompts and sends the results back to the server. The server receives these responses and integrates them. Specifically, it combines the results of the drama script generation AI with the results of the scenery drawing AI.
[1738] Specific examples
[1739] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[1740] Improving product quality through call and response
[1741] The server improves the quality of the generated results by retransmitting the results of the first-stage generation AI to the next-stage generation AI. This process is called call and response, and allows results to be passed between multiple generation AIs.
[1742] Specific examples
[1743] The server sends the results of the drama script generation AI to the scenery drawing AI, which then uses that information to generate new results. For example, the scenery drawing AI might return a result such as "depict the scenery at the moment the main character casts a spell in more detail."
[1744] Integration and delivery of the final product
[1745] The server aggregates the final product and sends it to the user's device, ensuring that the user receives a consistent, high-quality product at once.
[1746] Specific examples
[1747] The server sends the integrated results to the user as a "detailed script for the scene where magic is used in the square" and a "scenery description of that scene."
[1748] Use of the Generative AI App Store
[1749] Users can search, select, purchase, and review new AI generators through the AI generator app store, and the server will recommend the best AI generators to users based on these ratings and reviews.
[1750] Specific examples
[1751] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[1752] As described above, the present invention provides a unified management platform that allows users to efficiently utilize multiple generative AIs and obtain high-quality, consistent results.
[1753] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1754] Step 1:
[1755] The user selects a project and enters a prompt. This information is entered by accessing the platform from the user's device. The user selects the prompt and the generation AI to use.
[1756] Specific actions
[1757] The user accesses the platform's web page using a browser or app. On the project selection screen, they select "Drama Script" and then enter "Please create a scene where the main character uses magic for the first time" on the prompt input screen. They also select "Drama Script Generation AI" and "Landscape Drawing AI" as the generation AIs to use.
[1758] input
[1759] The prompt entered by the user, "Please create a scene where the main character uses magic for the first time," and the selection information of the generated AI.
[1760] output
[1761] The user's input data is temporarily saved in the device's internal storage.
[1762] Step 2:
[1763] The terminal transmits the prompt received from the user and information about the selected generated AI to the server.
[1764] Specific actions
[1765] When the user clicks the "Submit" button, the device converts the prompt and AI selection information into JSON format and sends it to the server.
[1766] input
[1767] User-entered prompts and generated AI selection information.
[1768] output
[1769] The data is sent to the server in JSON format.
[1770] Step 3:
[1771] The server converts the prompts received from the user into a format suitable for each generation AI and sends them to multiple generation AIs.
[1772] Specific actions
[1773] The server uses a natural language processing module to analyze the prompt and convert it into an input format for each generation AI. For example, for the "drama script generation AI," it converts it into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," and for the "scenery drawing AI," it converts it into "Describe the scene at the moment the magical light is emitted in the square."
[1774] input
[1775] The prompt and generated AI selection information received by the server in JSON format.
[1776] output
[1777] The prompt is converted into a format suitable for the generation AI and sent to the corresponding generation AI.
[1778] Step 4:
[1779] The generation AI generates artifacts based on prompts.
[1780] Specific actions
[1781] Each AI generator uses its internal algorithm to create a result based on the prompts it receives. For example, the "drama script generator AI" generates a scenario, and the "landscape drawing AI" generates an illustration.
[1782] input
[1783] The prompt converted into a format suitable for generative AI.
[1784] output
[1785] The artifacts (scripts, illustrations, etc.) are generated and sent back to the server.
[1786] Step 5:
[1787] The server receives and integrates the products from each generation AI.
[1788] Specific actions
[1789] The server receives the response data from each generation AI and combines them into a single integrated product, for example, integrating a scene from a drama script with the corresponding landscape description.
[1790] input
[1791] The products received from each generating AI.
[1792] output
[1793] Integrated artifacts (e.g., "A detailed script for the magic scene in the square" and "A landscape description of the scene").
[1794] Step 6:
[1795] The server manages the call and response between the generation AIs and retransmits the results from the first generation AI to the next generation AI in order to improve the quality of the results.
[1796] Specific actions
[1797] The server retransmits the results of the drama script generation AI to the scenery drawing AI, which then generates new products based on that information.
[1798] input
[1799] Results generated by the first-stage generation AI.
[1800] output
[1801] Improved production from next-stage generation AI.
[1802] Step 7:
[1803] The server aggregates the final product and sends it to the user's device.
[1804] Specific actions
[1805] The server integrates the products obtained from all the generation AIs, formats them as the final product, and provides it to the user.
[1806] input
[1807] A product integrated from multiple generation AIs.
[1808] output
[1809] The final product is sent to the user's terminal.
[1810] Step 8:
[1811] Users search, select, purchase, and review new generative AI through the generative AI app store.
[1812] Specific actions
[1813] Users access the app store, search for and purchase a generative AI, and write a review. The server analyzes these ratings and reviews and recommends the best generative AI for the user.
[1814] input
[1815] User ratings and reviews.
[1816] output
[1817] Recommendation information from the server.
[1818] (Application example 1)
[1819] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1820] In conventional design generation systems using generative AI models, it was difficult to effectively integrate multiple generative AI models and easily create high-quality 3D designs and product introduction videos for virtual stores. Furthermore, collaboration between generative AI models was insufficient, preventing sufficient improvement in the quality of the results. Furthermore, users lacked appropriate information when searching for and selecting new generative AI models, making it difficult to find the optimal model.
[1821] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1822] In this invention, the server includes means for a user to input a prompt sentence and select multiple generative AI models, means for the server to convert the prompt sentence received from the user into a format suitable for each generative AI model and send it to the multiple generative AI models, means for the generative AI models to generate products based on the prompt sentence, means for the server to receive and integrate the products from each generative AI model, and means for the operator of the virtual store to generate high-quality 3D designs and product introduction videos. This allows users to effectively use multiple generative AI models and obtain high-quality products at once, making it possible to generate high-quality content in the virtual store.
[1823] "User device" means an electronic device operated by a user, including a smartphone, smart glasses, a head-mounted display, etc.
[1824] A "generative AI model" refers to an artificial intelligence algorithm that automatically creates a specific product based on a prompt entered by a user.
[1825] A "prompt" is a text expression that succinctly summarizes the input information and instructions that a user provides to a generative AI model.
[1826] "Products" are the deliverables that the generative AI model creates based on the prompt, including 3D designs and product introduction videos.
[1827] A "server" is a central processing unit that runs on a network and manages the exchange of data between the user's device and the generative AI model.
[1828] A "virtual store" is a virtual store that displays and sells products and services online.
[1829] "Call and response between generative AI models" is a technique in which generative AI models work together to repeatedly respond, improving the accuracy and quality of the products.
[1830] A "Generative AI Model App Store" is an online platform where users can search for, select, purchase, and review new generative AI models.
[1831] A system for implementing the present invention includes a user terminal, a server, and multiple generative AI models. The user terminal is an electronic device such as a smartphone, smart glasses, or a head-mounted display, and provides an interface for a user to input a prompt sentence and select a generative AI model.
[1832] Hardware and software used
[1833] Hardware: Smartphones (e.g., iPhone, Samsung Galaxy), smart glasses (e.g., Google Glass), head-mounted displays (e.g., Meta Quest)
[1834] Software: iOS / Android app, web server, generative AI model (e.g., OpenAI's ChatGPT, DALL-E)
[1835] System Operation
[1836] 1. User's device:
[1837] Users access the application from their device, select a specific project, enter a prompt, and can select the generative AI model they want to use.
[1838] example:
[1839] I want to create a 3D model of a new product, and I need high-resolution images as well as a promotional video.
[1840] At this time, the user can select product design AI, video generation AI, or landscape drawing AI.
[1841] 2. Server prompt translation:
[1842] The server converts the prompt received from the user into a format suitable for each generative AI model, and then sends the converted prompt to each generative AI model.
[1843] 3. How generative AI models work:
[1844] The generative AI model creates a specified product based on the received prompt. For example, a product design AI generates a 3D model, a video generation AI generates a product introduction video, and a landscape drawing AI generates a background design.
[1845] 4. Server integration process:
[1846] The server receives the results from each generative AI model and integrates them. The integrated results are then sent to the user's device. At this time, a call-and-response function between the generative AI models is also activated, and the results of multiple generative AI models are sequentially linked to improve the quality of the results.
[1847] 5. User Interface:
[1848] The user can again review the product through the application, provide feedback on the product if necessary, and enter further prompts.
[1849] 6. Generative AI model app store:
[1850] Users can use the generative AI model app store to search for, select, purchase, and review new generative AI models, and the server has the function of recommending generative AI models based on ratings and reviews.
[1851] In this way, users can efficiently utilize a variety of generative AI models to easily generate high-quality 3D designs and product introduction videos for virtual stores.
[1852] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1853] Step 1:
[1854] A user accesses the application from their device, selects a specific project, and enters a prompt. The user selects the generative AI model to use, such as product design AI, video generation AI, or landscape drawing AI. The prompt entered might be something like, "I would like to create a 3D model of a new product. I also need a promotional video along with high-resolution images."
[1855] Step 2:
[1856] The terminal transmits the input prompt sentence and information about the selected generative AI model to the server. At this time, the input data from the terminal includes the prompt sentence and the ID information of the generative AI model.
[1857] Step 3:
[1858] The server analyzes the prompt received from the user and converts it into a format suitable for each generative AI model. For example, a prompt such as "I would like to create a 3D model of a new product" is converted into a format suitable for a product design AI to generate a high-resolution 3D model, and the part "We also need an introductory video" is converted into a format suitable for a video generation AI. The converted prompt is then sent to the generative AI model.
[1859] Step 4:
[1860] The generative AI models (product design AI, video generation AI, and landscape drawing AI) create the specified product based on the received prompt. The product design AI generates a high-resolution 3D model, the video generation AI generates a video introducing the product, and the landscape drawing AI generates a background design. Each generative AI model uses its internal algorithm to calculate data based on the prompt.
[1861] Step 5:
[1862] Each generative AI model completes the product and returns the data to the server. For example, a product design AI sends 3D model data, a video generation AI sends a video file, and a landscape drawing AI sends a background image.
[1863] Step 6:
[1864] The server integrates the results received from each generative AI model. It uses an integration algorithm to optimally combine the data, combining the 3D model data, video files, and background images into a single piece of content. During this integration process, a call-and-response function between the generative AI models is utilized to adjust the results to ensure higher quality.
[1865] Step 7:
[1866] The server transmits the integrated product to the user's terminal, which displays the received product and an interface for providing feedback if necessary.
[1867] Step 8:
[1868] When a user uses the generative AI model app store, the user can search, select, purchase, and review new generative AI models, and the server recommends the most suitable generative AI models based on the user's ratings and reviews.
[1869] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1870] The present invention provides a system that allows users to efficiently use multiple AI generators and further improves the quality of the output by recognizing the user's emotional state. The following describes in detail how the present invention is implemented.
[1871] User Interface
[1872] The process begins when a user accesses the platform from their device and logs in. The user selects a project on the interface screen, fills in prompts, and selects the generative AI to use. This interface works in conjunction with an emotion engine that also recognizes the user's emotions.
[1873] example:
[1874] The user inputs "I want to create the first act of a drama script" into the device and selects the "drama script generation AI." They also select the "landscape drawing AI" to draw the scenery. The emotion engine then analyzes the user's voice and facial expressions and recognizes them as "excited."
[1875] Sending prompts and selecting generation AI
[1876] The device sends the prompts entered by the user, information on the selected generation AI, and emotional data obtained from the emotion engine to the server, which then adjusts the prompts according to the user's emotional state.
[1877] example:
[1878] The prompt sent by the user, "Draw a scene where the main character uses magic for the first time," is received by the server, and the emotion engine detects the "excited" state, so the prompt is adjusted slightly.
[1879] Prompt conversion and generation AI transmission
[1880] The server converts the received prompts and emotion data into a format suitable for each generation AI and sends it to each generation AI. Data obtained from the emotion engine is also applied, so the generated content is adjusted to match the user's emotion.
[1881] example:
[1882] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," adds emotional data, and sends it to the drama script generation AI.
[1883] Response of the generating AI and reception of the product
[1884] The server receives responses from each AI generator. The AI generator creates a product based on each prompt and returns it to the server. The server then integrates these products and provides them to the user in the most optimal form.
[1885] example:
[1886] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[1887] Call and response between generative AI
[1888] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the results. This call-and-response process allows the generation AIs to cooperate with each other. Data from the emotion engine is also continuously used.
[1889] example:
[1890] The results of the drama script generation AI are sent to the scenery drawing AI, which then generates a new product based on that scene.
[1891] Integration and delivery of the final product
[1892] The server integrates the results from multiple generative AIs and provides the final result to the user, allowing users to receive high-quality results at once. The result is also further adjusted based on emotional data.
[1893] example:
[1894] The server sends the integrated results to the user as a detailed script for the scene in which the protagonist uses magic in the square and a description of the scenery of that scene. The description of the scene is adjusted to be even more dramatic depending on the user's level of excitement.
[1895] Use of the Generative AI App Store
[1896] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs based on the user's needs. This process also utilizes data from the emotion engine to recommend Generative AIs that are tailored to the user's current emotional state.
[1897] example:
[1898] Users select and install a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project." Users who are in an "excited" state will be suggested generative AIs that will further increase their excitement.
[1899] In this way, the present invention provides a centralized management platform for users to efficiently utilize multiple generative AIs and further improve the quality of the output using an emotion engine.
[1900] The processing flow will be explained below.
[1901] Step 1:
[1902] The user accesses the platform from their device and logs in. Login authentication is performed and the user's personal information is confirmed.
[1903] Specific behavior:
[1904] The user enters their ID and password on the login screen.
[1905] The device sends the login information to the server.
[1906] The server refers to the user database and performs authentication.
[1907] If the authentication is successful, the user will be taken to the dashboard screen.
[1908] Step 2:
[1909] The user selects a project, enters prompts, and selects each generated AI. The emotion engine then analyzes the user's emotional state.
[1910] Specific behavior:
[1911] On the device interface screen, select the desired project from the project list.
[1912] The user enters a prompt and selects the generated AI to use from a drop-down menu or list.
[1913] An emotion engine built into or connected to the device collects and analyzes emotional data from the user's voice and facial expressions.
[1914] The device sends prompts, the generated AI's selection information, and emotion data to the server.
[1915] Step 3:
[1916] The server receives the sent prompt, the selected generation AI information, and the emotion data from the emotion engine, and converts the prompt into a format suitable for each generation AI based on that information.
[1917] Specific behavior:
[1918] The server parses the prompt and converts it into a format to send to the generating AI.
[1919] Prepare to use emotion data obtained from the emotion engine to adjust the content of prompts and productions.
[1920] Create the necessary requests according to the API endpoint of each generation AI.
[1921] Step 4:
[1922] The server sends the converted prompt to each generation AI, requesting it to generate a product.
[1923] Specific behavior:
[1924] The server sends an HTTP request to the generated AI's API endpoint.
[1925] Each generation AI receives a prompt and creates the specified product.
[1926] Step 5:
[1927] Each generation AI creates a product and returns a response to the server, which receives these products.
[1928] Specific behavior:
[1929] Each generation AI generates a product based on a prompt.
[1930] The server receives the HTTP response from the generation AI and obtains the generated data.
[1931] Step 6:
[1932] The products received by the server are further adjusted and integrated through call and response between each generation AI.
[1933] Specific behavior:
[1934] The server sends the result from the first stage generation AI to the next stage generation AI as a prompt.
[1935] The next generation AI creates a new product and returns it to the server.
[1936] Repeat this process as necessary.
[1937] At each stage, the emotion engine data is used to adjust the output.
[1938] Step 7:
[1939] The final integrated product is sent to the user's device by the server, where it is further adjusted based on the emotion data.
[1940] Specific behavior:
[1941] The server converts the aggregated product into the appropriate format (text, images, other media).
[1942] The server further adjusts the product based on data from the emotion engine and sends it to the user's device.
[1943] Step 8:
[1944] Users can review the results they receive, make corrections and provide feedback as needed, and use the Generative AI App Store to find new Generative AI.
[1945] Specific behavior:
[1946] The user checks the output on the terminal.
[1947] Make corrections as needed and send feedback to the server.
[1948] The device accesses the generative AI app store, and the user searches for, selects, and purchases new generative AI.
[1949] The server recommends a generative AI based on user feedback.
[1950] Example 2
[1951] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1952] Conventional generative AI systems make it difficult for users to efficiently use multiple generative AIs, and have limited means for improving the consistency and quality of the results. Furthermore, they often fail to meet user expectations because they are unable to adjust the results to take into account the user's emotional state. The present invention aims to solve these problems and provide high-quality results that respond to the user's emotional state.
[1953] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for converting prompts and emotion data received from the user into a format suitable for each generation AI and transmitting the converted data to the multiple generation AIs, means for the generation AIs to generate products based on the prompts and emotion data, and means for the server to receive and integrate the products from the generation AIs. This makes it possible to adjust prompts based on the user's emotional state and consistently provide high-quality products.
[1954] A "user" is a person or institution that utilizes the system to enter prompts and select a generating AI.
[1955] "Terminal" means an electronic device used by a user to access the system and provide input information.
[1956] The "server" is a central processing unit that processes information received from users, converts it into a format suitable for the generative AI model, integrates the generated results, and provides them to users.
[1957] "Generative AI" is an artificial intelligence model that creates a specified artifact based on user-entered prompts.
[1958] A "prompt" is an instruction that the user enters to specify the desired product for the generation AI.
[1959] A "product" is an output result that the generation AI generates based on a prompt.
[1960] "Emotional data" refers to data that indicates the user's current emotional state and is used to adjust the product.
[1961] The "emotion engine" is a function that analyzes the user's voice, facial expressions, etc. to generate emotional data.
[1962] "Call and response" is a process in which the results obtained from the first-stage generation AI are retransmitted to the next-stage generation AI to improve the quality of the product.
[1963] The "Generative AI App Store" is an online platform where users can search for, select, purchase, and review new generative AI.
[1964] "Recommendation" means that the server suggests a generative AI that is suitable for the user based on ratings and reviews.
[1965] "Format conversion" is the process by which the server converts the prompt and emotion data received from the user into a format suitable for the generative AI.
[1966] "Integration" refers to the server combining the products obtained from each generation AI into one.
[1967] The system of the present invention is a platform that allows users to efficiently utilize multiple generation AIs and improve the quality of the products based on the user's emotional state. The components and specific functions of this system will be described below.
[1968] User Interface
[1969] Users access the platform from their own device and log in. Next, they select a project and enter a prompt for the AI generator. The user then selects the AI generator they want to use. At this time, an emotion engine is connected to the device, which analyzes and recognizes the user's emotional state in real time.
[1970] Example prompt sentence:
[1971] "I want to create the first act of a drama script."
[1972] Select "Drama script generation AI" and also select "Landscape drawing AI" for landscape drawing.
[1973] The emotion engine analyzes the user's voice and facial expressions to detect an "excited" state.
[1974] Sending prompt information and emotion data
[1975] The device sends the prompt text entered by the user, the selected generation AI information, and the emotion data obtained from the emotion engine to the server, which receives this information and prepares it for processing.
[1976] Prompt adjustment and transformation
[1977] The server analyzes the received prompt and emotion data, adjusts the prompt content as needed, and then converts it into the optimal format for each generative AI model before sending it.
[1978] Examples:
[1979] The server converts the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail," adds emotional data about excitement, and sends it to the drama script generation AI.
[1980] Receiving and integrating generative AI responses
[1981] The server receives responses from each AI generator. Each AI generator creates a product based on the prompt and emotion data and sends it back to the server. The server then integrates these products and provides them to the user in the most optimal form.
[1982] Examples:
[1983] The server receives and integrates the response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience is filled with surprise and cheers" and the response from the landscape drawing AI: "The square is lush with greenery, and the magical light shines blue."
[1984] Call and response between generative AI
[1985] The server retransmits the results of the first-stage generation AI to the next-stage generation AI to further improve the quality of the product. In this process, collaboration between the generation AIs is realized, and emotion data is also continuously used.
[1986] Examples:
[1987] The server sends the results of the drama script generation AI to the scenery drawing AI, which then generates a new product based on that scene.
[1988] Providing the final product
[1989] The server integrates the results from multiple generative AIs and provides the final result to the user. The result is further adjusted based on the user's emotional data, allowing the user to receive high-quality results all at once.
[1990] Examples:
[1991] The server sends the integrated results to the user as a "detailed script for the scene in which the protagonist uses magic in the square" and a "scenery description of that scene."
[1992] The depiction of the scene is adjusted to become more dramatic depending on the user's level of excitement.
[1993] Use of the Generative AI App Store
[1994] Users can access the Generative AI App Store to search, select, purchase, and review new Generative AIs. The server then recommends Generative AIs that meet the user's needs. Data from the emotion engine is also utilized to suggest Generative AIs that are suited to the user's current emotional state.
[1995] Examples:
[1996] The user selects and installs a highly rated generative AI from the app store, and the server recommends it, saying, "This generative AI is suitable for your project."
[1997] For users who are excited, generative AI will be suggested to further increase their excitement.
[1998] In this way, the present invention provides a unified platform that allows users to efficiently utilize multiple generative AI models and provide the highest quality products based on emotion data.
[1999] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2000] Step 1:
[2001] A user accesses the platform from their device and logs in. At this time, the user enters authentication information (username, password) and is granted access to the system. If the entered authentication information is correct, the server starts a session and displays the user interface (UI). This UI includes a project selection screen and a generation AI selection screen.
[2002] Input: Username, Password
[2003] Output: Session starts, UI is displayed
[2004] Specific behavior:
[2005] The user enters their credentials on the login page and clicks the submit button.
[2006] The server verifies the authentication information and, if correct, displays the dashboard screen.
[2007] Step 2:
[2008] The user selects a project on the interface screen and inputs a prompt for the AI to generate. The user also selects the AI to use. The device is connected to an emotion engine that analyzes and recognizes the user's emotional state in real time.
[2009] Input: Project selection, prompt text, AI generation selection
[2010] Output: Project information, prompt information, emotion data
[2011] Specific behavior:
[2012] The user enters "I want to create the first act of a drama script" and selects "Drama script generation AI."
[2013] The emotion engine scans the user's facial expressions to detect "excitement."
[2014] Step 3:
[2015] The device sends the prompt text entered by the user, information on the selected generation AI, and emotion data obtained from the emotion engine to the server, which receives this information and prepares to adjust the prompt text.
[2016] Input: prompt, generated AI information, emotion data
[2017] Output: Send data to the server
[2018] Specific behavior:
[2019] The terminal sends project information, prompts and "excitement state" data to the server.
[2020] Step 4:
[2021] The server analyzes the received prompt and emotion data, adjusts the prompt as needed, and then converts it into a format optimal for each generative AI model and sends it.
[2022] Input: prompt sentence, emotion data
[2023] Output: Adjusted prompt sentence, data sent to the generation AI
[2024] Specific behavior:
[2025] The server translates the prompt into "Scene 1: The protagonist uses magic for the first time in the square. Describe the scene and emotions in detail."
[2026] Data adjusted to a format suitable for the generative AI model is sent to each generative AI.
[2027] Step 5:
[2028] The generation AI generates a product based on the prompt sentence and emotion data and sends it back to the server. The server receives responses from each generation AI and integrates the products.
[2029] Input: Adjusted prompt sentence, emotion data
[2030] Output: Artifacts (e.g. text, images)
[2031] Specific behavior:
[2032] Generative AI generates text and images based on the prompts it receives.
[2033] The server receives and integrates the products returned by the generation AI.
[2034] example:
[2035] Response from the drama script generation AI: "The protagonist unleashes a magical light in the square. The audience erupts in amazement and cheers."
[2036] The landscape drawing AI responds, "The square is lush and green, and the magic light is blue."
[2037] Step 6:
[2038] The server retransmits the results of the first-stage AI generation to the second-stage AI generation to further improve the quality of the product. Emotion data is also continuously used. This results in a product of improved quality.
[2039] Input: Generation results from the first stage generation AI, emotion data
[2040] Output: Refined product from next stage generation AI
[2041] Specific behavior:
[2042] The server sends the results of the first stage drama script generation AI to the scenery drawing AI.
[2043] The landscape drawing AI generates new products based on the received data and sends them back.
[2044] Step 7:
[2045] The server integrates the results from multiple generative AIs and provides the final result to the user. Emotional data is also taken into account, allowing the user to receive the results in the most optimal way.
[2046] Input: Products from each generation AI, emotion data
[2047] Output: The integrated final product
[2048] Specific behavior:
[2049] The server integrates the results of each generation AI to create the final scene script and scenery description.
[2050] The scene depiction is adjusted to be even more dramatic depending on the user's level of excitement.
[2051] Step 8:
[2052] Users access the generative AI app store to search, select, purchase, and review new generative AIs. The server then recommends generative AIs that meet the user's needs, utilizing emotional data to make suggestions.
[2053] Input: User searches, selections, purchase information, ratings and reviews, sentiment data
[2054] Output: Recommendations and generative AI
[2055] Specific behavior:
[2056] A user searches for a new generative AI in the app store, checks reviews, and then purchases it.
[2057] The server recommends, "This generative AI is suitable for your project."
[2058] Through these specific steps and operations, the system of the present invention provides a high-performance and user-friendly generative AI platform.
[2059] (Application example 2)
[2060] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2061] Conventional generative AI systems lack the ability to customize their outputs based on the user's emotional state. This can result in outputs that do not match the user's expectations or emotional state, resulting in a poor user experience. Furthermore, there is also the issue of insufficient collaboration between generative AIs, resulting in a decline in the quality of the outputs. Furthermore, it is difficult to easily search for and select new generative AIs, preventing users from effectively utilizing a wide range of generative AIs.
[2062] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2063] In this invention, the server includes means for recognizing the emotional state of the user and adjusting the quality of the product based on the emotional data, an emotion engine for analyzing the emotional data, means for managing call and response between the generation AIs and retransmitting results from the first-stage generation AI to the next-stage generation AI to improve the quality of the product, means for the user to search, select, purchase, and review new generation AIs through a generation AI app store, and means for the server to recommend generation AI models based on the ratings and reviews. This makes it possible to provide high-quality products tailored to the user's emotional state, strengthens cooperation between generation AIs, and makes it easier to use new generation AIs.
[2064] "User" refers to the person using the generative AI system to input prompts.
[2065] A "prompt" refers to a sentence in the form of specific instructions or questions that are input to the generating AI.
[2066] "Generative AI" refers to an artificial intelligence model that creates artifacts based on prompts from the user.
[2067] "Server" refers to the computer system that receives prompts from the user, converts them into a format suitable for the generating AI, and then integrates and provides the resulting results to the user.
[2068] "Product" refers to the output data that the generative AI creates based on the prompts.
[2069] "Emotional state" refers to the user's current emotional or mental state, as recognized by the emotion engine.
[2070] "Emotion Engine" refers to a software or hardware system that analyzes a user's emotional state and provides that data.
[2071] "Call and response" refers to the collaborative process of retransmitting the product obtained from the first-stage generation AI to the next-stage generation AI to improve the quality of the product.
[2072] "Generative AI App Store" refers to an online platform where users can search for, select, purchase, and review new generative AI.
[2073] "Recommendation" refers to the act of suggesting the most suitable generative AI based on user ratings and reviews.
[2074] "User's device" refers to the keywords for smartphones, computers, tablets, etc. that users use to operate the generated AI system.
[2075] The present invention provides a system that provides high-quality products by efficiently utilizing multiple generation AIs while taking into account the emotional state of the user. Detailed embodiments of the present invention will be described below.
[2076] System Overview
[2077] The system of the present invention includes means for a user to input a prompt and select a generation AI, means for a server to convert the prompt into a suitable format and send it to the generation AI, means for the generation AI to generate a product based on the prompt, means for integrating the product and providing it to the user, and means for using an emotion engine to analyze the user's emotional state and adjust the product based thereon.
[2078] Hardware and software used
[2079] Hardware
[2080] User devices: smartphones, computers, tablets, etc.
[2081] Camera module: For capturing images to analyze the user's facial expressions
[2082] software
[2083] Emotion recognition engine: Software that analyzes the user's emotional state (e.g., the Hugging Face transformer library)
[2084] Generative AI models: Artificial intelligence models that create artifacts based on prompts (e.g., GPT-3)
[2085] Review analysis engine: Software that analyzes reviews based on artifacts
[2086] System Operation
[2087] 1. User Interface
[2088] Users access the interface from a device such as a smartphone or computer and log in. They then perform a search, select a project, fill in the prompts, and can also select a generative AI model.
[2089] 2. Acquiring Emotion Data
[2090] The camera module of the user's device is used to capture facial images, which are then analyzed by an emotion recognition engine to understand the user's emotional state, for example, determining whether the user is in an "excited" state.
[2091] 3. Sending prompts and selecting generation AI
[2092] The prompts and emotion data entered by the user are sent to the server, which converts them into a suitable format and sends them to the selected generative AI model.
[2093] 4. Product generation and preparation
[2094] The generative AI creates products based on prompts and adjusts the content based on emotional data. For example, if a user is in an excited state and enters the prompt, "I'm looking for a black leather jacket. Can you recommend any products?", the generative AI will generate product recommendations that correspond to the user's excited state.
[2095] 5. Consolidating and Presenting Results
[2096] The server integrates the results from each AI generator and provides it to the user, who then adjusts the results to match the user's emotional state.
[2097] 6. Use of Generative AI App Store
[2098] Users can search, select, purchase, and review new generative AI models through the generative AI app store, and the server will recommend the most suitable generative AI based on user ratings and reviews.
[2099] Specific examples
[2100] If a user types a prompt like "I'm looking for a black leather jacket. Can you recommend some?", and the user's emotional state is recognized as "excited," the following happens:
[2101] 1. Prompt and emotion data acquisition
[2102] Prompt: "I'm looking for a black leather jacket. Can you recommend one?"
[2103] Emotional state: "Excited"
[2104] 2. Product generation
[2105] The generative AI model generates detailed product descriptions and recommendations that match the user's excitement level, such as, "This black leather jacket is trending this season and has received very high ratings."
[2106] In this way, the present invention can provide high quality products based on the user's emotional state.
[2107] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2108] Step 1:
[2109] A user uses a terminal to access the interface and log in. The user selects a project and fills in the prompts. Then, the user selects multiple generative AI models.
[2110] Input: User login information, project selection, prompt text, and selection of generative AI model
[2111] Output: Information about the input prompt and the selected generative AI model
[2112] Specific behavior: For example, the user enters the prompt, "I'm looking for a black leather jacket. Can you recommend some?" and selects the relevant generative AI model.
[2113] Step 2:
[2114] The device's camera module is used to capture the user's facial expression, and the image data is sent to the emotion recognition engine for analysis.
[2115] Input: User's face image
[2116] Output: Parsed user emotional state data
[2117] Specific operation: A facial image is captured with a camera, sent to an emotion recognition engine, and the user's emotional state (e.g., excitement) is analyzed.
[2118] Step 3:
[2119] The prompt sentence and emotional state data are sent to the server, which receives the prompt sentence and emotional state data and converts the prompt sentence into a format suitable for each generative AI model.
[2120] Input: prompt sentence, emotional state data
[2121] Output: A prompt in a format suitable for a generative AI model
[2122] What happens: The server translates the prompt to something like, "I'm looking for a black leather jacket. Can you recommend one? [Emotion: Excited]."
[2123] Step 4:
[2124] The server sends the converted prompt sentence to each generative AI model, which creates a product based on the prompt sentence and returns it to the server.
[2125] Input: A prompt in a format suitable for a generative AI model
[2126] Output: The product of the generative AI model
[2127] Specific operation: The generative AI model generates product information and returns a result such as "This black leather jacket is a trendy item this season and has received very high ratings" to the server.
[2128] Step 5:
[2129] The server receives the outputs from each generative AI model, adjusts them based on the emotional state data, and integrates them.
[2130] Input: Product from generative AI model, emotional state data
[2131] Output: The integrated product
[2132] What it does: The server integrates the generated product details and recommendations based on them, dynamically adjusting them to fit the sentiment.
[2133] Step 6:
[2134] The server sends the integrated product to the user's terminal, and the user checks the received product.
[2135] Input: The integrated product
[2136] Output: The artifact provided to the user
[2137] Specific operation: The server sends a message to the user's device such as "Recommended product: This black leather jacket is a trendy item this season and has received very high ratings."
[2138] Step 7:
[2139] Users search, select, purchase, and review new generative AI models through the generative AI app store, and the server recommends generative AI models based on ratings and reviews.
[2140] Input: User ratings and reviews
[2141] Output: A generative AI model recommended to the user
[2142] Specific operation: The user searches for new generative AI models in the generative AI app store, purchases them, and reviews them. The server then uses this information to suggest the optimal generative AI model.
[2143] Through the above processing steps, this system realizes the operation of an advanced generative AI model that takes into account the user's emotional state, providing high-quality results.
[2144] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2145] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2146] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2147] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2148] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2149] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2150] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2151] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2152] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2153] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2154] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2155] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2156] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2157] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2158] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2159] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2160] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2161] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2162] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2163] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2164] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2165] The following is further disclosed regarding the above embodiment.
[2166] (Claim 1)
[2167] A means for the user to enter prompts and select from multiple generated AIs;
[2168] A means for the server to convert prompts received from users into a format suitable for each generation AI and send them to multiple generation AIs;
[2169] a means for the generation AI to generate a product based on the prompt;
[2170] A means for the server to receive and integrate the products from each generation AI;
[2171] means for the server to transmit the integrated product to the user's terminal;
[2172] A system including:
[2173] (Claim 2)
[2174] The system of claim 1, wherein the server manages call and response between the generation AIs and includes means for retransmitting results from a first-stage generation AI to a next-stage generation AI in order to improve the quality of the product.
[2175] (Claim 3)
[2176] A means for users to search, select, purchase, and review new generative AI through the generative AI app store;
[2177] 10. The system of claim 1, wherein the server includes means for recommending generative AI based on ratings and reviews.
[2178] "Example 1"
[2179] (Claim 1)
[2180] A means for the user to enter prompts and select from multiple generated AIs;
[2181] A means for transmitting information about the prompt received by the terminal from the user and the selected generation AI to the server;
[2182] A means for the server to convert prompts received from users into a format suitable for each ge...
Claims
1. A means for the user to enter prompts and select from multiple generated AIs; A means for the server to convert prompts received from users into a format suitable for each generation AI and send them to multiple generation AIs; a means for the generation AI to generate a product based on the prompt; A means for the server to receive and integrate the products from each generation AI; means for the server to transmit the integrated product to the user's terminal; A system including:
2. The system of claim 1, wherein the server manages call and response between the generation AIs and includes means for retransmitting results from the first-stage generation AI to the next-stage generation AI in order to improve the quality of the product.
3. A means for users to search, select, purchase, and review new generative AI through the generative AI app store; The system of claim 1 , wherein the server includes means for recommending a generative AI based on ratings and reviews.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A