system
The system addresses the limitations of current AI agents by allowing users to customize visual and audio settings and transfer data, ensuring consistent and emotionally responsive AI experiences.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-12
- Publication Date
- 2026-06-24
AI Technical Summary
Current AI agent systems lack visual and voice personalization, flexibility in customization, and smooth data transfer, failing to meet user preferences and provide consistent experiences.
A system that allows users to select from multiple templates, customize visual and audio settings, generate personalized AI agents, and transfer usage data to new agents, while enabling purchases through a marketplace.
Enables users to enjoy highly personalized AI experiences with customizable agents that maintain consistency and adapt to individual preferences and emotional states, enhancing user satisfaction.
Smart Images

Figure 2026103461000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] With the rapid growth of the AI agent market, users are demanding more personalized experiences and natural conversations. However, current systems have insufficient visual and voice personalization, so there is a need to improve the user experience. In addition, users are seeking means to easily customize agents according to their preferences, but conventional systems lack flexibility and have problems with smooth data transfer after purchase and utilization.
Means for Solving the Problems
[0005] This invention expands the scope of personalization by providing users with multiple selectable templates and a means to customize the visuals and audio based on them. Furthermore, it realizes the specific experiences users desire by providing a means to generate and deliver customized agents. In addition, it addresses individual user needs by providing a marketplace where users can purchase and download additional content to their devices, while simultaneously supporting the smooth use of agents by leveraging past usage history by providing a means to transfer usage data from existing agents to new agents.
[0006] A "template" refers to a combination of visuals and sounds that a user can select as the initial settings for an AI agent.
[0007] "Customization" refers to the act of adjusting the visual and audio details of a template selected by the user, and changing the settings to suit their preferences.
[0008] An "agent" refers to an AI program that can interact with users and respond to their instructions and questions.
[0009] "Generation" refers to the process of creating a new agent based on customized settings.
[0010] A "marketplace" refers to an online platform where users can purchase agent-related products such as additional visual and audio content.
[0011] "Downloading" refers to the process of transferring and saving purchased or acquired digital data to a user's device via the internet.
[0012] "Data transfer" refers to the process of migrating data such as user usage history and configuration information from an existing agent to a newly customized agent. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention provides a system that allows users to customize AI agents according to their preferences and enjoy a personalized experience. Users can select an agent from a variety of templates and customize its appearance and voice down to the smallest detail. The system's operation is described below.
[0035] First, the user accesses the platform using their device and creates an account. During this process, they enter the necessary authentication information, and upon successful login, they are redirected to the home screen.
[0036] Next, the server provides the user with several AI agent templates. The user browses through them and selects their preferred template. The selected template displays visual and voice settings menus, which the user uses to customize it.
[0037] Once customization is complete, the server collects the configuration information and generates a customized AI agent for that specific user. This generated agent is then provided to the user's device and ready for interaction.
[0038] Furthermore, by utilizing the marketplace, users can purchase additional visuals and audio created by other creators. The server manages this content and makes the corresponding data available for download to the user's device once the user's purchase is confirmed.
[0039] Purchased agents and customized data are transferred to new agents using the user's past usage data, ensuring a consistent user experience for continued use.
[0040] For example, a user can select a "female character with a pop music style," adjust the pitch and speed of her voice, and change the color of her appearance to their preferred color. This agent is then generated with the selected custom voice and appearance and runs on the user's device.
[0041] In this way, the present invention enables users to obtain an experience tailored to their individual needs through an AI agent. Furthermore, it allows for deeper levels of personalization by leveraging unique customization options.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] The user's device accesses the platform and displays an account creation screen. Here, the user enters the required information, such as name, email address, and password, to create an account. The server receives the entered data, verifies its accuracy, and then stores it in the database.
[0045] Step 2:
[0046] Authentication is performed by accessing the login screen using the user's device and entering the registered email address and password. The server compares the entered information with the database information, and if successful, redirects the user to the home screen.
[0047] Step 3:
[0048] The server provides the user with a list of multiple AI agent templates on the home screen. The user's device receives this list and displays a template selection screen. The user can then choose their preferred template.
[0049] Step 4:
[0050] Once a template is selected, the user's device displays a visual and audio customization menu. The user can customize the details using tools such as sliders and color selectors. The server records the user's changes in real time and saves the configuration data.
[0051] Step 5:
[0052] Once the user has finished customizing their settings, the server generates a unique AI agent based on that configuration data. This generated agent is sent from the server to the user's device, which the user can then download and begin interacting with.
[0053] Step 6:
[0054] Users are directed to the marketplace where they can browse additional visual and audio content. Once a user decides to purchase, the server interacts with an external payment processing service to complete the transaction. After the purchase is confirmed, a download link is provided, making the content available on the user's device.
[0055] Step 7:
[0056] The server analyzes past agent usage data and transfers it to a new, customized agent. The user's device receives this data and verifies that consistent settings and conversation history are maintained on the new agent.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] Traditional intelligent agent customization had the problem of not being able to adequately address the individual needs of users. Furthermore, the complexity of the operation when making various custom settings made it difficult for users to enjoy intuitive and flexible customization. In addition, existing systems did not allow for smooth retrieval of additional information or transfer of history, resulting in an inconsistent user experience.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, the server includes means for providing visual and auditory information as multiple information sources selectable by the user, means for individually configuring the selected information sources via the user's terminal, and means for the server to generate an intelligent agent corresponding to the user's individual configuration information. This allows the user to intuitively select information sources from multiple options, easily configure them individually, and enjoy an intelligent agent tailored to their individual needs.
[0062] "Information base" refers to a diverse set of visual and auditory features that users can select, and it is the foundation that determines the appearance and voice of an intelligent agent.
[0063] An "intelligent agent" is an interactive program generated by the server based on the user's individual settings, providing functions tailored to the user's specific needs.
[0064] A "trading marketplace" is an online platform where users can purchase additional information and extensions, offering them a variety of options.
[0065] "History transfer" is the process of applying past information created or used by the user to a new intelligent agent, ensuring a consistent user experience.
[0066] "Financial transactions through external services" refers to payment and post-processing activities conducted through services provided by third parties in connection with purchases on trading markets.
[0067] This invention provides a system that allows users to customize intelligent agents according to their preferences and enjoy a personalized experience. Users access the system using a network-connected terminal and utilize the services through that terminal.
[0068] The server retrieves multiple visual and auditory information sources from the database and provides them to the user. The user can browse these information sources through the terminal interface and select the desired ones. For the selected information sources, the user uses the terminal to configure individual settings and set the visual and auditory details.
[0069] Examples of devices used include personal computers, tablets, and smartphones. The server uses a generative AI model to generate an intelligent agent based on the user's settings. This process integrates selected visual and auditory information and performs data editing and analysis within the program.
[0070] Furthermore, users can obtain additional information and enhancements through the trading market provided by the server. The content selected by the user can be downloaded to the device and applied to the intelligent agent. During this process, the server references the user's past usage history and transfers necessary historical information to the new intelligent agent to provide a consistent experience.
[0071] As a concrete example, a user can select a "female character with a pop music style," adjust the pitch and speed of her voice, and change the color of her appearance to their preferred color. This agent is then generated with the selected custom voice and appearance and runs on the user's device.
[0072] Examples of prompt messages include: "Provide specific customization options based on the template selected by the user," and "Describe the process by which the server collects customization settings and generates a new AI agent."
[0073] This system allows users to intuitively customize and enjoy a personalized intelligent agent tailored to their specific needs.
[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0075] Step 1:
[0076] The user accesses the system using their device, creates an account, and logs in. As input, the user provides an email address and password. The server then verifies the authentication information and identifies the user. As output, the home screen is displayed on the user's device.
[0077] Step 2:
[0078] The server provides the user with multiple visual and auditory information sources. The user browses the information sources via an interface on the terminal and selects the desired information sources. The selected information sources are sent to the server as input. Based on this, the server processes the selected information sources and returns them to the user. The information sources selected by the user are displayed on the terminal as output.
[0079] Step 3:
[0080] The user customizes the selected information base. Specifically, they adjust the character's appearance and voice using the settings screen on their device. The input includes the customization parameters specified by the user. The server receives these parameters, analyzes the data using a generative AI model, and generates an intelligent agent. As output, a preview of the customized agent is displayed on the user's device.
[0081] Step 4:
[0082] The server formally generates an intelligent agent based on the customized information and provides it to the user's terminal. The user's final customized data is used as input. The server performs data calculations and sends the generation results to the user's terminal. As output, the fully generated intelligent agent becomes available for use on the terminal.
[0083] Step 5:
[0084] Users can obtain additional information through the trading market. Here, the server provides market information upon user request. Input includes the user's trading requests and payment information. The server processes this information and, if approved, provides purchase data. Output is the additional information downloaded to the user's terminal and reflected in the intelligent agent.
[0085] Step 6:
[0086] The server performs a procedure to inherit past agent usage data. The server receives the user's historical data as input. The server analyzes this data and applies it to the latest agent. The output is an intelligent agent running on the user's terminal, guaranteeing a consistent user experience.
[0087] (Application Example 1)
[0088] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0089] In customizing information processing agents, it is necessary for users to have a personalized experience down to the smallest detail according to their preferences, and to provide a friendly and attractive machine by adjusting the robot's personality and appearance. To achieve this goal, an easily accessible electronic marketplace and a smooth purchasing process are required.
[0090] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0091] In this invention, the server includes means for providing a user with a set of templates to choose from, means for the user to adjust visual and auditory features based on the templates, and means for generating and providing the adjusted information processing agent to the user. This makes it possible to easily realize the customized experience desired by the user and provide a superior user experience.
[0092] A "user" is an individual or legal entity that uses this system to customize information processing agents and enjoy a personalized experience.
[0093] A "template" is a design template that includes pre-defined visual and auditory features, serving as the basis for users to customize information processing agents.
[0094] "Visual features" refer to attributes related to the appearance of an information processing agent, including elements such as color and design.
[0095] "Speech characteristics" refer to the attributes of the voice emitted by an information processing agent, including tone, speed, and intonation.
[0096] An "information processing agent" is a program or machine device that has user-customized visual and auditory characteristics and is designed to perform a specific task.
[0097] An "electronic marketplace" is an online platform where users can purchase additional information and data, and it provides content to expand the customization of agents.
[0098] "Usage data" refers to the history and configuration information of when an existing information processing agent was used, and this information is carried over to the new agent.
[0099] "Individuality" refers to the uniqueness and expressiveness of a robot, and its characteristics are adjusted in appearance and behavior according to the user's preferences.
[0100] In an embodiment of this invention, a user first accesses the system using a dedicated interface and creates their own account. After creating an account, the user can select an information processing agent from a variety of templates. The server provides these templates and allows customization of the selected visual and auditory features. Customization involves setting auditory features using a speech synthesis library (e.g., pyttsx3) and displaying visual features using a graphics library (e.g., pygame).
[0101] The server generates a customized agent and delivers it to the user's terminal. The platform on the terminal provides the ability to purchase additional information data through an electronic marketplace, and the purchased information is designed to be transferred to the terminal immediately. This allows the user to give the customized robot a friendly voice and personality.
[0102] As a concrete example, the AI agent of a cleaning robot can be personalized with a Hawaiian theme and act as a friendly character. Users can select specific voices, such as, "Aloha! Have a great day, I'll start by cleaning the kitchen!" Such customization makes using the robot more enjoyable.
[0103] An example of a prompt to input into the generating AI model is, "Explain how to personalize the AI agent for a cleaning robot with a Hawaiian theme and create a user-friendly character."
[0104] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0105] Step 1:
[0106] The user accesses the system using their terminal and creates an account. The user's authentication information (e.g., username, password) is required as input. The server receives this information, stores it in the database, and generates a new user account. A message confirming successful account creation is displayed on the user's terminal as output.
[0107] Step 2:
[0108] The server presents the user with templates for multiple information processing agents. User account information is required as input. The server verifies the user's access permissions and prepares a list of templates. The list of templates is displayed on the user's terminal as output.
[0109] Step 3:
[0110] The user selects a template and customizes its visual and auditory features. The input requires the user's selected template and customizations (e.g., color scheme, voice tone). The device records this information and displays a realistic preview in real time. The output displays the specific results customized by the user.
[0111] Step 4:
[0112] The server generates a new information processing agent based on the customized settings and makes it available for download to the user's terminal. User customization information is required as input. The server builds the agent and generates compiled data based on this information. As output, a download link is displayed on the user's terminal.
[0113] Step 5:
[0114] The user purchases additional information data from an electronic marketplace. The input requires the user's payment information and selected product information. The server integrates with an external payment service and approves the purchase upon successful payment. The output is a purchase success notification displayed on the user's device, and the product data becomes available for download.
[0115] Step 6:
[0116] The server transfers usage data from the existing agent to the new agent. The server requires usage history data from the existing agent as input. It analyzes this data and synchronizes it with the new agent. The output is a consistent user experience on the new agent, and a notification of the transfer completion is displayed on the user's device.
[0117] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0118] This invention provides a platform that allows users to customize AI agents according to their preferences, and by combining it with an emotion engine, it constructs a system that enables responses based on the user's emotions. The operation of this system is described below.
[0119] First, the user accesses the platform using their device, registers an account, and logs in. The server verifies the user's credentials and guides them to the home screen. Next, the user can select their preferred AI agent template from several options and customize its visuals and voice in detail on their device. The server receives the user's customization information in real time and generates a unique AI agent.
[0120] The generated AI agent is not only provided to the user's device, but also incorporates an emotion engine. During the interaction between the user and the agent, the device senses the user's voice and facial expressions. The emotion engine analyzes these inputs and estimates the user's emotional state. The server adjusts the agent's response according to the estimated emotion. For example, if the user seems sad, the agent will be configured to use a more friendly voice and write encouraging words.
[0121] Furthermore, the system includes a feature that allows users to transfer usage data from existing agents to new user customizations, providing a consistent user experience. Additional content from the marketplace is managed by the server, and users can purchase and download new visuals and sounds to their devices.
[0122] As a concrete example, consider a scenario where a user selects a "business assistant," sets a formal voice, and chooses visuals with somewhat subdued colors. If the emotion engine analyzes that the user is experiencing stress, the agent will be set to offer advice to alleviate stress or to converse in a relaxed tone.
[0123] Such a multi-functional AI agent not only executes user commands but also enables personalized dialogue that is sensitive to the user's emotions. This invention enriches the user experience and enables more natural interaction.
[0124] The following describes the processing flow.
[0125] Step 1:
[0126] The user's device accesses the platform and displays the account registration screen. The user enters their name, email address, and password to create an account. The server receives the input data, checks for duplicates with existing data, and then saves it to the database.
[0127] Step 2:
[0128] The user accesses the login screen using their device and enters their registered email address and password. The server compares this information with the database, and if authentication is successful, it displays the home screen to the user.
[0129] Step 3:
[0130] The server provides the user with a list of AI agent templates on the home screen. The user's device displays this list, and the user can select a template that suits their preferences. The selected template information is then sent to the server.
[0131] Step 4:
[0132] Once a template is selected, the user's device provides visual and audio customization menus. Users can adjust colors, voice tone, and other attributes using sliders and pull-down menus to create their own personalized settings. The server analyzes and stores this customization data in real time.
[0133] Step 5:
[0134] Once customization is complete, the server generates a unique AI agent based on the collected data. The generated agent is then sent from the server to the user's device, where the user can use it.
[0135] Step 6:
[0136] An emotion engine is embedded in the user's device. This emotion engine captures and analyzes the user's voice and facial expressions through the microphone and camera. The server receives this analysis and estimates the user's emotions.
[0137] Step 7:
[0138] Based on the estimated emotional information, the server adjusts the AI agent's response. For example, if it senses that the user is tired, the agent is programmed to send an encouraging message in a cheerful voice. The terminal then presents this response to the user.
[0139] Step 8:
[0140] When a user decides to purchase new visuals or audio through the marketplace, the device displays a purchase screen. The server integrates with an external payment system to complete the purchase process. Once the purchase is confirmed, a download link is provided to the device, allowing the user to access the new content.
[0141] Step 9:
[0142] The server begins the process of transferring existing usage data to the new customized agent. This data migration ensures that users can enjoy a consistent experience with the new agent. The terminal notifies the user when this transfer is complete.
[0143] (Example 2)
[0144] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0145] Traditional AI agent systems have had problems with their inability to flexibly respond to changes in user emotions. Furthermore, their customization interfaces are cumbersome, and changes are not reflected in real time. Additionally, the acquisition of additional content and payment procedures are cumbersome for users.
[0146] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0147] In this invention, the server includes means for providing multiple styles that the user can select, means for the user to modify visual and auditory elements based on the style, and means for performing emotion analysis and adjusting the function's response based on the user's emotional state. This allows the user to easily customize a flexible AI agent that is attuned to their emotions and seamlessly acquire additional content and complete payment procedures.
[0148] "Style" refers to a template that allows users to select the basic structure and characteristics of an AI agent.
[0149] "Visual elements" refer to the components related to the appearance, colors, and design of an AI agent.
[0150] "Acoustic elements" refer to the components related to the quality, tone, and intonation of the AI agent's voice.
[0151] "Functionality" refers to the capabilities and behaviors of an AI agent that have been modified and generated by the user.
[0152] The "marketplace" refers to a virtual trading platform where users can acquire additional content for their AI agents.
[0153] "Device" refers to a device used by a user to interact with an AI agent.
[0154] "Usage information" refers to the data and history accumulated when a user uses the AI agent.
[0155] "Emotion analysis" refers to a technology that infers a user's emotional state from their voice and facial expressions.
[0156] This invention is implemented using a terminal used by the user and a server that processes data. First, the user accesses the platform from the terminal and logs in with their account. The server verifies the user's authentication information and executes the login process.
[0157] Subsequently, the user can select the most suitable AI agent style from several options on the platform. Based on the selected style, an interface is provided on the device to make detailed modifications to the visual and auditory elements. The server receives this modification information and uses the generated AI model to build an AI agent tailored to the user's customizations.
[0158] The AI agent has emotion analysis capabilities, and the device collects voice and posture data during user interaction. This data is analyzed by a server, and the agent's response is adjusted based on the user's emotional state. This technology utilizes natural language processing and image recognition algorithms. For example, if the user has a sad expression, the agent will be configured to use more friendly and encouraging words.
[0159] Users can also acquire additional visual and auditory elements from the marketplace via their devices. The server can transfer purchased materials to the user's device and reflect the new materials without affecting existing agents. Furthermore, a data transfer function ensures that user usage information is carried over to the new agent, maintaining a consistent user experience.
[0160] As a concrete example, consider a scenario where a user selects "Professional Assistant" and sets a formal voice and simple design. In this case, if emotion analysis reveals that the user is experiencing stress, the agent will adjust to provide advice to help them relax.
[0161] An example of a prompt message for a generative AI model would be: "Based on the style of the AI agent selected by the user, analyze the emotions in real time and generate an appropriate response."
[0162] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0163] Step 1:
[0164] Users use their devices to access the platform, enter their account information, and log in.
[0165] Upon receiving the entered authentication information, the server accesses the database to verify its validity. If the information is correct, the user is redirected to the home screen. Specifically, if the user forgets their password, a password reset option is provided.
[0166] Step 2:
[0167] The user selects from several AI agent styles on the home screen.
[0168] The format selected by the terminal is sent to the server. The server returns information about that format to the user and provides data to display a visual preview of the options.
[0169] In terms of specific operation, the user is prompted to decide on the optimal format while viewing a preview.
[0170] Step 3:
[0171] Based on the selected style, users can make detailed modifications to visual and auditory elements on their device. This involves adjusting colors and sound tones using sliders and selectors.
[0172] Correction input is sent from the terminal to the server. The server uses a generation AI model to process the correction data and performs calculations to generate the corrected AI agent.
[0173] The generated data is sent back to the user's terminal, and the corrected results are previewed.
[0174] Step 4:
[0175] The user begins interacting with the generated AI agent.
[0176] The device uses sensors to record the user's voice and facial expressions, and then sends that data to a server.
[0177] The server performs data analysis, uses an emotion analysis engine to identify the user's emotional state, and processes the data so that a generative AI model can generate an appropriate response.
[0178] In terms of specific actions, if the user appears sad, the agent will respond in a gentle tone, saying something like, "Please let me know if there's anything I can help you with."
[0179] Step 5:
[0180] Users can use the marketplace to purchase additional visual and audio elements and download them to their devices.
[0181] Purchase operations are sent from the terminal to the server, and payment is completed by linking with an external payment service.
[0182] The server receives purchase information and processes the data to transfer the necessary content to the user's device.
[0183] The device instantly deploys downloaded content and reflects it in the AI agent. Specifically, the user can check their purchase history and be presented with an option to try new features immediately.
[0184] (Application Example 2)
[0185] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0186] In recent years, digital agents have become increasingly widespread, but many of these agents are unable to empathize with users' emotions during interactions, remaining limited to simply providing information or executing commands. Therefore, there is a need for systems that can provide more natural and personalized communication, especially for the elderly and users who require emotional support. Furthermore, the customizability of agents and the diversity of their content are limited, failing to adequately address the diverse preferences and needs of users.
[0187] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0188] In this invention, the server includes means for providing a user-selectable set of templates, means for the user to customize visual and auditory elements based on the templates, and means for generating and providing a customized agent to the user. This makes it possible to analyze the user's emotional state and adjust the agent's response based on the analysis results. This allows for personalized communication that is attentive to the user's emotions, thereby improving user satisfaction. Furthermore, by utilizing the market to purchase and make downloadable diverse content, a system with scalability to meet diverse user needs is realized.
[0189] A "template" is a set of settings that forms the basis for users to select the appearance and behavior of an agent.
[0190] "Visual" refers to elements related to the appearance and graphics of the user's digital agent.
[0191] "Acoustics" refers to elements related to the voices and sound effects that digital agents emit to users.
[0192] An "agent" is a digital interface that interacts with users, providing information and taking actions based on instructions.
[0193] A "marketplace" is a platform where users can purchase additional content, and where content is provided to expand the user's agent.
[0194] A "server" is a computing system that generates and personalizes digital agents in response to user requests and provides the results to the user.
[0195] "Emotional state" refers to the emotions a user is currently experiencing and serves as foundational information for adjusting the agent's response.
[0196] To implement this invention, a system is required for generating and operating digital agents for users. The server provides the user with multiple selectable templates, from which the user can customize the agent's visual and auditory properties. The server receives the user's selection in real time and generates the customized agent. In doing so, a database is used to consider the user's past usage data, ensuring a consistent user experience.
[0197] The terminal is expected to be a wearable device such as smart glasses. On the terminal, the agent interacts with the user. During this interaction, the terminal's camera and microphone are used to collect the user's facial expressions and voice data. This data is sent to a server, where an emotion engine analyzes the user's emotional state. Based on the analysis results, the agent adjusts its voice and tone of voice, enabling it to respond in a way that is empathetic to the user's stress and anxiety. For the implementation of the emotion engine, it is recommended to use services such as Google Cloud's Vision API or Amazon Transcribe.
[0198] As a concrete example, imagine an elderly person exchanging daily information with an agent through smart glasses, and the displayed agent gently speaking to the user according to their emotions and offering stress-reducing advice. In this case, the agent might gently encourage the user by saying, "I recommend you relax today. Let's take a few deep breaths."
[0199] An example of a prompt message would be, "Based on the elderly person's facial expressions and voice data, continue the conversation in a relaxed tone and provide advice to help them have a comfortable day." This allows for the use of a generative AI model that adjusts the agent's response.
[0200] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0201] Step 1:
[0202] The user accesses the server using a terminal and selects their preferred digital agent from several templates. The user ID and template ID are provided as input, which the server receives. Based on the template, visual and auditory options are generated and presented to the user. As output, the visual and auditory customization options are displayed on the user's terminal.
[0203] Step 2:
[0204] The user customizes the visual and auditory settings on their device and sends this information to the server. The server processes this customization data in real time as input. The final customization data is saved as output, forming the basis for generating the customized agent. A generative AI model is then used to generate prompts that create an agent based on the customizations.
[0205] Step 3:
[0206] The server utilizes a generative AI model to generate an agent based on stored customized data and delivers it to the user's terminal. Inputs include customized data and generated prompt messages. Output is a digital agent with matching visual and auditory elements displayed on the user's terminal.
[0207] Step 4:
[0208] The device continuously captures the user's facial expressions and voice using its camera and microphone, and sends this data to the server. The input includes voice and image data, which the server analyzes using an emotion engine. The output is an estimated result of the user's emotional state, and this data is used to adjust the agent's response. Specifically, data processing and emotion analysis are performed using Google Cloud's Vision API and Amazon Transcribe.
[0209] Step 5:
[0210] The server adjusts the agent's voice and tone of voice based on the inferred emotional state to tailor its response to the user. The emotion analysis results are included as input, and the adjusted voice response from the agent is played back on the user's device as output. Specifically, the response will be in a relaxed tone to reduce the user's anxiety and stress.
[0211] Step 6:
[0212] The system records the agent's responses to the user and saves them as data to help improve future interactions. Inputs include the agent's responses and user feedback. Outputs include accumulated reference data to improve the consistency of responses and the personalized experience in subsequent interactions.
[0213] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0214] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0215] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0216] [Second Embodiment]
[0217] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0218] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0219] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0220] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0221] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0222] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0223] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0224] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0225] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0226] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0227] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0228] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0229] This invention provides a system that allows users to customize AI agents according to their preferences and enjoy a personalized experience. Users can select an agent from a variety of templates and customize its appearance and voice down to the smallest detail. The system's operation is described below.
[0230] First, the user accesses the platform using their device and creates an account. During this process, they enter the necessary authentication information, and upon successful login, they are redirected to the home screen.
[0231] Next, the server provides the user with several AI agent templates. The user browses through them and selects their preferred template. The selected template displays visual and voice settings menus, which the user uses to customize it.
[0232] Once customization is complete, the server collects the configuration information and generates a customized AI agent for that specific user. This generated agent is then provided to the user's device and ready for interaction.
[0233] Furthermore, by utilizing the marketplace, users can purchase additional visuals and audio created by other creators. The server manages this content and makes the corresponding data available for download to the user's device once the user's purchase is confirmed.
[0234] Purchased agents and customized data are transferred to new agents using the user's past usage data, ensuring a consistent user experience for continued use.
[0235] For example, a user can select a "female character with a pop music style," adjust the pitch and speed of her voice, and change the color of her appearance to their preferred color. This agent is then generated with the selected custom voice and appearance and runs on the user's device.
[0236] In this way, the present invention enables users to obtain an experience tailored to their individual needs through an AI agent. Furthermore, it allows for deeper levels of personalization by leveraging unique customization options.
[0237] The following describes the processing flow.
[0238] Step 1:
[0239] The user's device accesses the platform and displays an account creation screen. Here, the user enters the required information, such as name, email address, and password, to create an account. The server receives the entered data, verifies its accuracy, and then stores it in the database.
[0240] Step 2:
[0241] Authentication is performed by accessing the login screen using the user's device and entering the registered email address and password. The server compares the entered information with the database information, and if successful, redirects the user to the home screen.
[0242] Step 3:
[0243] The server provides the user with a list of multiple AI agent templates on the home screen. The user's device receives this list and displays a template selection screen. The user can then choose their preferred template.
[0244] Step 4:
[0245] Once a template is selected, the user's device displays a visual and audio customization menu. The user can customize the details using tools such as sliders and color selectors. The server records the user's changes in real time and saves the configuration data.
[0246] Step 5:
[0247] Once the user has finished customizing their settings, the server generates a unique AI agent based on that configuration data. This generated agent is sent from the server to the user's device, which the user can then download and begin interacting with.
[0248] Step 6:
[0249] Users are directed to the marketplace where they can browse additional visual and audio content. Once a user decides to purchase, the server interacts with an external payment processing service to complete the transaction. After the purchase is confirmed, a download link is provided, making the content available on the user's device.
[0250] Step 7:
[0251] The server analyzes past agent usage data and transfers it to a new, customized agent. The user's device receives this data and verifies that consistent settings and conversation history are maintained on the new agent.
[0252] (Example 1)
[0253] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0254] Traditional intelligent agent customization had the problem of not being able to adequately address the individual needs of users. Furthermore, the complexity of the operation when making various custom settings made it difficult for users to enjoy intuitive and flexible customization. In addition, existing systems did not allow for smooth retrieval of additional information or transfer of history, resulting in an inconsistent user experience.
[0255] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0256] In this invention, the server includes means for providing visual and auditory information as multiple information sources selectable by the user, means for individually configuring the selected information sources via the user's terminal, and means for the server to generate an intelligent agent corresponding to the user's individual configuration information. This allows the user to intuitively select information sources from multiple options, easily configure them individually, and enjoy an intelligent agent tailored to their individual needs.
[0257] "Information base" refers to a diverse set of visual and auditory features that users can select, and it is the foundation that determines the appearance and voice of an intelligent agent.
[0258] An "intelligent agent" is an interactive program generated by the server based on the user's individual settings, providing functions tailored to the user's specific needs.
[0259] A "trading marketplace" is an online platform where users can purchase additional information and extensions, offering them a variety of options.
[0260] "History transfer" is the process of applying past information created or used by the user to a new intelligent agent, ensuring a consistent user experience.
[0261] "Financial transactions through external services" refers to payment and post-processing activities conducted through services provided by third parties in connection with purchases on trading markets.
[0262] This invention provides a system that allows users to customize intelligent agents according to their preferences and enjoy a personalized experience. Users access the system using a network-connected terminal and utilize the services through that terminal.
[0263] The server retrieves multiple visual and auditory information sources from the database and provides them to the user. The user can browse these information sources through the terminal interface and select the desired ones. For the selected information sources, the user uses the terminal to configure individual settings and set the visual and auditory details.
[0264] Examples of devices used include personal computers, tablets, and smartphones. The server uses a generative AI model to generate an intelligent agent based on the user's settings. This process integrates selected visual and auditory information and performs data editing and analysis within the program.
[0265] Furthermore, users can obtain additional information and enhancements through the trading market provided by the server. The content selected by the user can be downloaded to the device and applied to the intelligent agent. During this process, the server references the user's past usage history and transfers necessary historical information to the new intelligent agent to provide a consistent experience.
[0266] As a concrete example, a user can select a "female character with a pop music style," adjust the pitch and speed of her voice, and change the color of her appearance to their preferred color. This agent is then generated with the selected custom voice and appearance and runs on the user's device.
[0267] Examples of prompt messages include: "Provide specific customization options based on the template selected by the user," and "Describe the process by which the server collects customization settings and generates a new AI agent."
[0268] This system allows users to intuitively customize and enjoy a personalized intelligent agent tailored to their specific needs.
[0269] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0270] Step 1:
[0271] The user accesses the system using their device, creates an account, and logs in. As input, the user provides an email address and password. The server then verifies the authentication information and identifies the user. As output, the home screen is displayed on the user's device.
[0272] Step 2:
[0273] The server provides the user with multiple visual and auditory information sources. The user browses the information sources via an interface on the terminal and selects the desired information sources. The selected information sources are sent to the server as input. Based on this, the server processes the selected information sources and returns them to the user. The information sources selected by the user are displayed on the terminal as output.
[0274] Step 3:
[0275] The user customizes the selected information base. Specifically, they adjust the character's appearance and voice using the settings screen on their device. The input includes the customization parameters specified by the user. The server receives these parameters, analyzes the data using a generative AI model, and generates an intelligent agent. As output, a preview of the customized agent is displayed on the user's device.
[0276] Step 4:
[0277] The server formally generates an intelligent agent based on the customized information and provides it to the user's terminal. The user's final customized data is used as input. The server performs data calculations and sends the generation results to the user's terminal. As output, the fully generated intelligent agent becomes available for use on the terminal.
[0278] Step 5:
[0279] The user can obtain additional information through the trading market. Here, the server provides market information according to the user's request. As input, the user's trading request and payment information are included. The server processes these and, if approved, provides purchase data. As output, the additional information is downloaded to the user's terminal and reflected in the intelligent agent.
[0280] Step 6:
[0281] The server executes a procedure to inherit past agent usage data. As input, the user's past data exists on the server. The server analyzes this and applies it to the latest agent. As output, an intelligent agent that guarantees a consistent user experience operates on the user's terminal.
[0282] (Application Example 1)
[0283] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0284] In the customization of the information processing agent, it is possible for the user to obtain a personalized experience down to the details according to their preferences, and further, by adjusting the personality and appearance of the robot, it is required to provide a friendly and attractive mechanical device. To achieve this goal, an easily accessible electronic market and a smooth purchasing procedure are necessary.
[0285] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0286] In this invention, the server includes means for providing a plurality of templates that can be selected by the user, means for the user to adjust visual features and audio features based on the templates, and means for generating and providing to the user an adjusted information processing agent. This makes it possible to easily realize a customized experience desired by the user and provide an excellent user experience.
[0287] The "user" is an individual or legal entity who uses this system to customize an information processing agent and enjoy an individual experience.
[0288] A "template" is a design template that serves as a basis when the user customizes an information processing agent and includes pre-prepared visual and audio features.
[0289] "Visual features" are attributes related to the appearance of an information processing agent and include elements such as color and design.
[0290] "Audio features" are attributes of the voice emitted by an information processing agent and include aspects such as tone, speed, and intonation of the voice.
[0291] An "information processing agent" is a program or mechanical device that has visual features and audio features customized by the user and is used to execute specific tasks.
[0292] An "electronic market" is an online platform where users can purchase additional information data and is a place that provides content for expanding the customization of agents.
[0293] "Usage data" is historical and setting information when an existing information processing agent is used and is information passed on to a new agent.
[0294] "Personality" refers to the uniqueness and expressiveness of a robot and is a characteristic whose appearance and behavior are adjusted according to the user's preferences.
[0295] In an embodiment of this invention, a user first accesses the system using a dedicated interface and creates their own account. After creating an account, the user can select an information processing agent from a variety of templates. The server provides these templates and allows customization of the selected visual and auditory features. Customization involves setting auditory features using a speech synthesis library (e.g., pyttsx3) and displaying visual features using a graphics library (e.g., pygame).
[0296] The server generates a customized agent and delivers it to the user's terminal. The platform on the terminal provides the ability to purchase additional information data through an electronic marketplace, and the purchased information is designed to be transferred to the terminal immediately. This allows the user to give the customized robot a friendly voice and personality.
[0297] As a concrete example, the AI agent of a cleaning robot can be personalized with a Hawaiian theme and act as a friendly character. Users can select specific voices, such as, "Aloha! Have a great day, I'll start by cleaning the kitchen!" Such customization makes using the robot more enjoyable.
[0298] An example of a prompt to input into the generating AI model is, "Explain how to personalize the AI agent for a cleaning robot with a Hawaiian theme and create a user-friendly character."
[0299] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0300] Step 1:
[0301] The user accesses the system using a terminal and creates an account. As input, the user's authentication information (e.g., username, password) is required. The server receives this information, saves it in the database, and generates a new user account. As output, a success message for account creation is displayed on the user's terminal.
[0302] Step 2:
[0303] The server presents the user with prototypes of multiple information processing agents. As input, the user's account information is required. The server checks the user's access rights and prepares a list of prototypes. As output, a list of prototypes is displayed on the user's terminal.
[0304] Step 3:
[0305] The user selects a prototype and customizes its visual and audio features. As input, the user's selected prototype and customization content (e.g., color, voice tone) are required. The terminal records this information and displays a realistic preview in real time. As output, the specific results customized by the user are displayed.
[0306] Step 4:
[0307] The server generates a new information processing agent based on the customized settings and makes it downloadable to the user's terminal. As input, the user's customization information is required. The server constructs an agent based on this information and generates compiled data. As output, a download link is displayed on the user's terminal.
[0308] Step 5:
[0309] 1] The user purchases additional information data from an electronic marketplace. The input requires the user's payment information and selected product information. The server integrates with an external payment service and approves the purchase upon successful payment. The output is a purchase success notification displayed on the user's device, and the product data becomes available for download.
[0310] Step 6:
[0311] The server transfers usage data from the existing agent to the new agent. The server requires usage history data from the existing agent as input. It analyzes this data and synchronizes it with the new agent. The output is a consistent user experience on the new agent, and a notification of the transfer completion is displayed on the user's device.
[0312] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0313] This invention provides a platform that allows users to customize AI agents according to their preferences, and by combining it with an emotion engine, it constructs a system that enables responses based on the user's emotions. The operation of this system is described below.
[0314] First, the user accesses the platform using their device, registers an account, and logs in. The server verifies the user's credentials and guides them to the home screen. Next, the user can select their preferred AI agent template from several options and customize its visuals and voice in detail on their device. The server receives the user's customization information in real time and generates a unique AI agent.
[0315] The generated AI agent is not only provided to the user's device, but also incorporates an emotion engine. During the interaction between the user and the agent, the device senses the user's voice and facial expressions. The emotion engine analyzes these inputs and estimates the user's emotional state. The server adjusts the agent's response according to the estimated emotion. For example, if the user seems sad, the agent will be configured to use a more friendly voice and write encouraging words.
[0316] Furthermore, the system includes a feature that allows users to transfer usage data from existing agents to new user customizations, providing a consistent user experience. Additional content from the marketplace is managed by the server, and users can purchase and download new visuals and sounds to their devices.
[0317] As a concrete example, consider a scenario where a user selects a "business assistant," sets a formal voice, and chooses visuals with somewhat subdued colors. If the emotion engine analyzes that the user is experiencing stress, the agent will be set to offer advice to alleviate stress or to converse in a relaxed tone.
[0318] Such a multi-functional AI agent not only executes user commands but also enables personalized dialogue that is sensitive to the user's emotions. This invention enriches the user experience and enables more natural interaction.
[0319] The following describes the processing flow.
[0320] Step 1:
[0321] The user's device accesses the platform and displays the account registration screen. The user enters their name, email address, and password to create an account. The server receives the input data, checks for duplicates with existing data, and then saves it to the database.
[0322] Step 2:
[0323] The user accesses the login screen using their device and enters their registered email address and password. The server compares this information with the database, and if authentication is successful, it displays the home screen to the user.
[0324] Step 3:
[0325] The server provides the user with a list of AI agent templates on the home screen. The user's device displays this list, and the user can select a template that suits their preferences. The selected template information is then sent to the server.
[0326] Step 4:
[0327] Once a template is selected, the user's device provides visual and audio customization menus. Users can adjust colors, voice tone, and other attributes using sliders and pull-down menus to create their own personalized settings. The server analyzes and stores this customization data in real time.
[0328] Step 5:
[0329] Once customization is complete, the server generates a unique AI agent based on the collected data. The generated agent is then sent from the server to the user's device, where the user can use it.
[0330] Step 6:
[0331] An emotion engine is embedded in the user's device. This emotion engine captures and analyzes the user's voice and facial expressions through the microphone and camera. The server receives this analysis and estimates the user's emotions.
[0332] Step 7:
[0333] Based on the estimated emotional information, the server adjusts the AI agent's response. For example, if it senses that the user is tired, the agent is programmed to send an encouraging message in a cheerful voice. The terminal then presents this response to the user.
[0334] Step 8:
[0335] When a user decides to purchase new visuals or audio through the marketplace, the device displays a purchase screen. The server integrates with an external payment system to complete the purchase process. Once the purchase is confirmed, a download link is provided to the device, allowing the user to access the new content.
[0336] Step 9:
[0337] The server begins the process of transferring existing usage data to the new customized agent. This data migration ensures that users can enjoy a consistent experience with the new agent. The terminal notifies the user when this transfer is complete.
[0338] (Example 2)
[0339] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0340] Traditional AI agent systems have had problems with their inability to flexibly respond to changes in user emotions. Furthermore, their customization interfaces are cumbersome, and changes are not reflected in real time. Additionally, the acquisition of additional content and payment procedures are cumbersome for users.
[0341] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0342] In this invention, the server includes means for providing multiple styles that the user can select, means for the user to modify visual and auditory elements based on the style, and means for performing emotion analysis and adjusting the function's response based on the user's emotional state. This allows the user to easily customize a flexible AI agent that is attuned to their emotions and seamlessly acquire additional content and complete payment procedures.
[0343] "Style" refers to a template that allows users to select the basic structure and characteristics of an AI agent.
[0344] "Visual elements" refer to the components related to the appearance, colors, and design of an AI agent.
[0345] "Acoustic elements" refer to the components related to the quality, tone, and intonation of the AI agent's voice.
[0346] "Functionality" refers to the capabilities and behaviors of an AI agent that have been modified and generated by the user.
[0347] The "marketplace" refers to a virtual trading platform where users can acquire additional content for their AI agents.
[0348] "Device" refers to a device used by a user to interact with an AI agent.
[0349] "Usage information" refers to the data and history accumulated when a user uses the AI agent.
[0350] "Emotion analysis" refers to a technology that infers a user's emotional state from their voice and facial expressions.
[0351] This invention is implemented using a terminal used by the user and a server that processes data. First, the user accesses the platform from the terminal and logs in with their account. The server verifies the user's authentication information and executes the login process.
[0352] Subsequently, the user can select the most suitable AI agent style from several options on the platform. Based on the selected style, an interface is provided on the device to make detailed modifications to the visual and auditory elements. The server receives this modification information and uses the generated AI model to build an AI agent tailored to the user's customizations.
[0353] The AI agent has emotion analysis capabilities, and the device collects voice and posture data during user interaction. This data is analyzed by a server, and the agent's response is adjusted based on the user's emotional state. This technology utilizes natural language processing and image recognition algorithms. For example, if the user has a sad expression, the agent will be configured to use more friendly and encouraging words.
[0354] Users can also acquire additional visual and auditory elements from the marketplace via their devices. The server can transfer purchased materials to the user's device and reflect the new materials without affecting existing agents. Furthermore, a data transfer function ensures that user usage information is carried over to the new agent, maintaining a consistent user experience.
[0355] As a concrete example, consider a scenario where a user selects "Professional Assistant" and sets a formal voice and simple design. In this case, if emotion analysis reveals that the user is experiencing stress, the agent will adjust to provide advice to help them relax.
[0356] An example of a prompt message for a generative AI model would be: "Based on the style of the AI agent selected by the user, analyze the emotions in real time and generate an appropriate response."
[0357] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0358] Step 1:
[0359] Users use their devices to access the platform, enter their account information, and log in.
[0360] Upon receiving the entered authentication information, the server accesses the database to verify its validity. If the information is correct, the user is redirected to the home screen. Specifically, if the user forgets their password, a password reset option is provided.
[0361] Step 2:
[0362] The user selects from several AI agent styles on the home screen.
[0363] The format selected by the terminal is sent to the server. The server returns information about that format to the user and provides data to display a visual preview of the options.
[0364] In terms of specific operation, the user is prompted to decide on the optimal format while viewing a preview.
[0365] Step 3:
[0366] Based on the selected style, users can make detailed modifications to visual and auditory elements on their device. This involves adjusting colors and sound tones using sliders and selectors.
[0367] Correction input is sent from the terminal to the server. The server uses a generation AI model to process the correction data and performs calculations to generate the corrected AI agent.
[0368] The generated data is sent back to the user's terminal, and the corrected results are previewed.
[0369] Step 4:
[0370] The user begins interacting with the generated AI agent.
[0371] The device uses sensors to record the user's voice and facial expressions, and then sends that data to a server.
[0372] The server performs data analysis, uses an emotion analysis engine to identify the user's emotional state, and processes the data so that a generative AI model can generate an appropriate response.
[0373] In terms of specific actions, if the user appears sad, the agent will respond in a gentle tone, saying something like, "Please let me know if there's anything I can help you with."
[0374] Step 5:
[0375] Users can use the marketplace to purchase additional visual and audio elements and download them to their devices.
[0376] Purchase operations are sent from the terminal to the server, and payment is completed by linking with an external payment service.
[0377] The server receives purchase information and processes the data to transfer the necessary content to the user's device.
[0378] The device instantly deploys downloaded content and reflects it in the AI agent. Specifically, the user can check their purchase history and be presented with an option to try new features immediately.
[0379] (Application Example 2)
[0380] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0381] In recent years, digital agents have become increasingly widespread, but many of these agents are unable to empathize with users' emotions during interactions, remaining limited to simply providing information or executing commands. Therefore, there is a need for systems that can provide more natural and personalized communication, especially for the elderly and users who require emotional support. Furthermore, the customizability of agents and the diversity of their content are limited, failing to adequately address the diverse preferences and needs of users.
[0382] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0383] In this invention, the server includes means for providing a user-selectable set of templates, means for the user to customize visual and auditory elements based on the templates, and means for generating and providing a customized agent to the user. This makes it possible to analyze the user's emotional state and adjust the agent's response based on the analysis results. This allows for personalized communication that is attentive to the user's emotions, thereby improving user satisfaction. Furthermore, by utilizing the market to purchase and make downloadable diverse content, a system with scalability to meet diverse user needs is realized.
[0384] A "template" is a set of settings that forms the basis for users to select the appearance and behavior of an agent.
[0385] "Visual" refers to elements related to the appearance and graphics of the user's digital agent.
[0386] "Acoustics" refers to elements related to the voices and sound effects that digital agents emit to users.
[0387] An "agent" is a digital interface that interacts with users, providing information and taking actions based on instructions.
[0388] A "marketplace" is a platform where users can purchase additional content, and where content is provided to expand the user's agent.
[0389] A "server" is a computing system that generates and personalizes digital agents in response to user requests and provides the results to the user.
[0390] "Emotional state" refers to the emotions a user is currently experiencing and serves as foundational information for adjusting the agent's response.
[0391] To implement this invention, a system is required for generating and operating digital agents for users. The server provides the user with multiple selectable templates, from which the user can customize the agent's visual and auditory properties. The server receives the user's selection in real time and generates the customized agent. In doing so, a database is used to consider the user's past usage data, ensuring a consistent user experience.
[0392] The terminal is expected to be a wearable device such as smart glasses. On the terminal, the agent interacts with the user. During this interaction, the terminal's camera and microphone are used to collect the user's facial expressions and voice data. This data is sent to a server, where an emotion engine analyzes the user's emotional state. Based on the analysis results, the agent adjusts its voice and tone of voice, enabling it to respond in a way that is empathetic to the user's stress and anxiety. For implementing the emotion engine, it is recommended to use services such as Google Cloud's Vision API or Amazon Transcribe.
[0393] As a concrete example, imagine an elderly person exchanging daily information with an agent through smart glasses, and the displayed agent gently speaking to the user according to their emotions and offering stress-reducing advice. In this case, the agent might gently encourage the user by saying, "I recommend you relax today. Let's take a few deep breaths."
[0394] An example of a prompt message would be, "Based on the elderly person's facial expressions and voice data, continue the conversation in a relaxed tone and provide advice to help them have a comfortable day." This allows for the use of a generative AI model that adjusts the agent's response.
[0395] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0396] Step 1:
[0397] The user accesses the server using a terminal and selects their preferred digital agent from several templates. The user ID and template ID are provided as input, which the server receives. Based on the template, visual and auditory options are generated and presented to the user. As output, the visual and auditory customization options are displayed on the user's terminal.
[0398] Step 2:
[0399] The user customizes the visual and auditory settings on their device and sends this information to the server. The server processes this customization data in real time as input. The final customization data is saved as output, forming the basis for generating the customized agent. A generative AI model is then used to generate prompts that create an agent based on the customizations.
[0400] Step 3:
[0401] The server utilizes a generative AI model to generate an agent based on stored customized data and delivers it to the user's terminal. Inputs include customized data and generated prompt messages. Output is a digital agent with matching visual and auditory elements displayed on the user's terminal.
[0402] Step 4:
[0403] The device continuously captures the user's facial expressions and voice using its camera and microphone, and sends this data to the server. The input includes voice and image data, which the server analyzes using an emotion engine. The output is an estimated result of the user's emotional state, and this data is used to adjust the agent's response. Specifically, data processing and emotion analysis are performed using Google Cloud's Vision API and Amazon Transcribe.
[0404] Step 5:
[0405] The server adjusts the agent's voice and tone of voice based on the inferred emotional state to tailor its response to the user. The emotion analysis results are included as input, and the adjusted voice response from the agent is played back on the user's device as output. Specifically, the response will be in a relaxed tone to reduce the user's anxiety and stress.
[0406] Step 6:
[0407] The system records the agent's responses to the user and saves them as data to help improve future interactions. Inputs include the agent's responses and user feedback. Outputs include accumulated reference data to improve the consistency of responses and the personalized experience in subsequent interactions.
[0408] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0409] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0410] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0411] [Third Embodiment]
[0412] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0413] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0414] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0415] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0416] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0417] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0418] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0419] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0420] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0421] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0422] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0423] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0424] This invention provides a system that allows users to customize AI agents according to their preferences and enjoy a personalized experience. Users can select an agent from a variety of templates and customize its appearance and voice down to the smallest detail. The system's operation is described below.
[0425] First, the user accesses the platform using their device and creates an account. During this process, they enter the necessary authentication information, and upon successful login, they are redirected to the home screen.
[0426] Next, the server provides the user with several AI agent templates. The user browses through them and selects their preferred template. The selected template displays visual and voice settings menus, which the user uses to customize it.
[0427] Once customization is complete, the server collects the configuration information and generates a customized AI agent for that specific user. This generated agent is then provided to the user's device and ready for interaction.
[0428] Furthermore, by utilizing the marketplace, users can purchase additional visuals and audio created by other creators. The server manages this content and makes the corresponding data available for download to the user's device once the user's purchase is confirmed.
[0429] Purchased agents and customized data are transferred to new agents using the user's past usage data, ensuring a consistent user experience for continued use.
[0430] For example, a user can select a "female character with a pop music style," adjust the pitch and speed of her voice, and change the color of her appearance to their preferred color. This agent is then generated with the selected custom voice and appearance and runs on the user's device.
[0431] In this way, the present invention enables users to obtain an experience tailored to their individual needs through an AI agent. Furthermore, it allows for deeper levels of personalization by leveraging unique customization options.
[0432] The following describes the processing flow.
[0433] Step 1:
[0434] The user's device accesses the platform and displays an account creation screen. Here, the user enters the required information, such as name, email address, and password, to create an account. The server receives the entered data, verifies its accuracy, and then stores it in the database.
[0435] Step 2:
[0436] Authentication is performed by accessing the login screen using the user's device and entering the registered email address and password. The server compares the entered information with the database information, and if successful, redirects the user to the home screen.
[0437] Step 3:
[0438] The server provides the user with a list of multiple AI agent templates on the home screen. The user's device receives this list and displays a template selection screen. The user can then choose their preferred template.
[0439] Step 4:
[0440] Once a template is selected, the user's device displays a visual and audio customization menu. The user can customize the details using tools such as sliders and color selectors. The server records the user's changes in real time and saves the configuration data.
[0441] Step 5:
[0442] Once the user has finished customizing their settings, the server generates a unique AI agent based on that configuration data. This generated agent is sent from the server to the user's device, which the user can then download and begin interacting with.
[0443] Step 6:
[0444] Users are directed to the marketplace where they can browse additional visual and audio content. Once a user decides to purchase, the server interacts with an external payment processing service to complete the transaction. After the purchase is confirmed, a download link is provided, making the content available on the user's device.
[0445] Step 7:
[0446] The server analyzes past agent usage data and transfers it to a new, customized agent. The user's device receives this data and verifies that consistent settings and conversation history are maintained on the new agent.
[0447] (Example 1)
[0448] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0449] Traditional intelligent agent customization had the problem of not being able to adequately address the individual needs of users. Furthermore, the complexity of the operation when making various custom settings made it difficult for users to enjoy intuitive and flexible customization. In addition, existing systems did not allow for smooth retrieval of additional information or transfer of history, resulting in an inconsistent user experience.
[0450] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0451] In this invention, the server includes means for providing visual and auditory information as multiple information sources selectable by the user, means for individually configuring the selected information sources via the user's terminal, and means for the server to generate an intelligent agent corresponding to the user's individual configuration information. This allows the user to intuitively select information sources from multiple options, easily configure them individually, and enjoy an intelligent agent tailored to their individual needs.
[0452] "Information base" refers to a diverse set of visual and auditory features that users can select, and it is the foundation that determines the appearance and voice of an intelligent agent.
[0453] An "intelligent agent" is an interactive program generated by the server based on the user's individual settings, providing functions tailored to the user's specific needs.
[0454] A "trading marketplace" is an online platform where users can purchase additional information and extensions, offering them a variety of options.
[0455] "History transfer" is the process of applying past information created or used by the user to a new intelligent agent, ensuring a consistent user experience.
[0456] "Financial transactions through external services" refers to payment and post-processing activities conducted through services provided by third parties in connection with purchases on trading markets.
[0457] This invention provides a system that allows users to customize intelligent agents according to their preferences and enjoy a personalized experience. Users access the system using a network-connected terminal and utilize the services through that terminal.
[0458] The server retrieves multiple visual and auditory information sources from the database and provides them to the user. The user can browse these information sources through the terminal interface and select the desired ones. For the selected information sources, the user uses the terminal to configure individual settings and set the visual and auditory details.
[0459] Examples of devices used include personal computers, tablets, and smartphones. The server uses a generative AI model to generate an intelligent agent based on the user's settings. This process integrates selected visual and auditory information and performs data editing and analysis within the program.
[0460] Furthermore, users can obtain additional information and enhancements through the trading market provided by the server. The content selected by the user can be downloaded to the device and applied to the intelligent agent. During this process, the server references the user's past usage history and transfers necessary historical information to the new intelligent agent to provide a consistent experience.
[0461] As a concrete example, a user can select a "female character with a pop music style," adjust the pitch and speed of her voice, and change the color of her appearance to their preferred color. This agent is then generated with the selected custom voice and appearance and runs on the user's device.
[0462] Examples of prompt messages include: "Provide specific customization options based on the template selected by the user," and "Describe the process by which the server collects customization settings and generates a new AI agent."
[0463] This system allows users to intuitively customize and enjoy a personalized intelligent agent tailored to their specific needs.
[0464] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0465] Step 1:
[0466] The user accesses the system using their device, creates an account, and logs in. As input, the user provides an email address and password. The server then verifies the authentication information and identifies the user. As output, the home screen is displayed on the user's device.
[0467] Step 2:
[0468] The server provides the user with multiple visual and auditory information sources. The user browses the information sources via an interface on the terminal and selects the desired information sources. The selected information sources are sent to the server as input. Based on this, the server processes the selected information sources and returns them to the user. The information sources selected by the user are displayed on the terminal as output.
[0469] Step 3:
[0470] The user customizes the selected information base. Specifically, they adjust the character's appearance and voice using the settings screen on their device. The input includes the customization parameters specified by the user. The server receives these parameters, analyzes the data using a generative AI model, and generates an intelligent agent. As output, a preview of the customized agent is displayed on the user's device.
[0471] Step 4:
[0472] The server formally generates an intelligent agent based on the customized information and provides it to the user's terminal. The user's final customized data is used as input. The server performs data calculations and sends the generation results to the user's terminal. As output, the fully generated intelligent agent becomes available for use on the terminal.
[0473] Step 5:
[0474] Users can obtain additional information through the trading market. Here, the server provides market information upon user request. Input includes the user's trading requests and payment information. The server processes this information and, if approved, provides purchase data. Output is the additional information downloaded to the user's terminal and reflected in the intelligent agent.
[0475] Step 6:
[0476] The server performs a procedure to inherit past agent usage data. The server receives the user's historical data as input. The server analyzes this data and applies it to the latest agent. The output is an intelligent agent running on the user's terminal, guaranteeing a consistent user experience.
[0477] (Application Example 1)
[0478] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0479] In customizing information processing agents, it is necessary for users to have a personalized experience down to the smallest detail according to their preferences, and to provide a friendly and attractive machine by adjusting the robot's personality and appearance. To achieve this goal, an easily accessible electronic marketplace and a smooth purchasing process are required.
[0480] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0481] In this invention, the server includes means for providing a user with a set of templates to choose from, means for the user to adjust visual and auditory features based on the templates, and means for generating and providing the adjusted information processing agent to the user. This makes it possible to easily realize the customized experience desired by the user and provide a superior user experience.
[0482] A "user" is an individual or legal entity that uses this system to customize information processing agents and enjoy a personalized experience.
[0483] A "template" is a design template that includes pre-defined visual and auditory features, serving as the basis for users to customize information processing agents.
[0484] "Visual features" refer to attributes related to the appearance of an information processing agent, including elements such as color and design.
[0485] "Speech characteristics" refer to the attributes of the voice emitted by an information processing agent, including tone, speed, and intonation.
[0486] An "information processing agent" is a program or machine device that has user-customized visual and auditory characteristics and is designed to perform a specific task.
[0487] An "electronic marketplace" is an online platform where users can purchase additional information and data, and it provides content to expand the customization of agents.
[0488] "Usage data" refers to the history and configuration information of when an existing information processing agent was used, and this information is carried over to the new agent.
[0489] "Individuality" refers to the uniqueness and expressiveness of a robot, and its characteristics are adjusted in appearance and behavior according to the user's preferences.
[0490] In an embodiment of this invention, a user first accesses the system using a dedicated interface and creates their own account. After creating an account, the user can select an information processing agent from a variety of templates. The server provides these templates and allows customization of the selected visual and auditory features. Customization involves setting auditory features using a speech synthesis library (e.g., pyttsx3) and displaying visual features using a graphics library (e.g., pygame).
[0491] The server generates a customized agent and delivers it to the user's terminal. The platform on the terminal provides the ability to purchase additional information data through an electronic marketplace, and the purchased information is designed to be transferred to the terminal immediately. This allows the user to give the customized robot a friendly voice and personality.
[0492] As a concrete example, the AI agent of a cleaning robot can be personalized with a Hawaiian theme and act as a friendly character. Users can select specific voices, such as, "Aloha! Have a great day, I'll start by cleaning the kitchen!" Such customization makes using the robot more enjoyable.
[0493] An example of a prompt to input into the generating AI model is, "Explain how to personalize the AI agent for a cleaning robot with a Hawaiian theme and create a user-friendly character."
[0494] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0495] Step 1:
[0496] The user accesses the system using their terminal and creates an account. The user's authentication information (e.g., username, password) is required as input. The server receives this information, stores it in the database, and generates a new user account. A message confirming successful account creation is displayed on the user's terminal as output.
[0497] Step 2:
[0498] The server presents the user with templates for multiple information processing agents. User account information is required as input. The server verifies the user's access permissions and prepares a list of templates. The list of templates is displayed on the user's terminal as output.
[0499] Step 3:
[0500] The user selects a template and customizes its visual and auditory features. The input requires the user's selected template and customizations (e.g., color scheme, voice tone). The device records this information and displays a realistic preview in real time. The output displays the specific results customized by the user.
[0501] Step 4:
[0502] The server generates a new information processing agent based on the customized settings and makes it available for download to the user's terminal. User customization information is required as input. The server builds the agent and generates compiled data based on this information. As output, a download link is displayed on the user's terminal.
[0503] Step 5:
[0504] The user purchases additional information data from an electronic marketplace. The input requires the user's payment information and selected product information. The server integrates with an external payment service and approves the purchase upon successful payment. The output is a purchase success notification displayed on the user's device, and the product data becomes available for download.
[0505] Step 6:
[0506] The server transfers usage data from the existing agent to the new agent. The server requires usage history data from the existing agent as input. It analyzes this data and synchronizes it with the new agent. The output is a consistent user experience on the new agent, and a notification of the transfer completion is displayed on the user's device.
[0507] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0508] This invention provides a platform that allows users to customize AI agents according to their preferences, and by combining it with an emotion engine, it constructs a system that enables responses based on the user's emotions. The operation of this system is described below.
[0509] First, the user accesses the platform using their device, registers an account, and logs in. The server verifies the user's credentials and guides them to the home screen. Next, the user can select their preferred AI agent template from several options and customize its visuals and voice in detail on their device. The server receives the user's customization information in real time and generates a unique AI agent.
[0510] The generated AI agent is not only provided to the user's device, but also incorporates an emotion engine. During the interaction between the user and the agent, the device senses the user's voice and facial expressions. The emotion engine analyzes these inputs and estimates the user's emotional state. The server adjusts the agent's response according to the estimated emotion. For example, if the user seems sad, the agent will be configured to use a more friendly voice and write encouraging words.
[0511] Furthermore, the system includes a feature that allows users to transfer usage data from existing agents to new user customizations, providing a consistent user experience. Additional content from the marketplace is managed by the server, and users can purchase and download new visuals and sounds to their devices.
[0512] As a concrete example, consider a scenario where a user selects a "business assistant," sets a formal voice, and chooses visuals with somewhat subdued colors. If the emotion engine analyzes that the user is experiencing stress, the agent will be set to offer advice to alleviate stress or to converse in a relaxed tone.
[0513] Such a multi-functional AI agent not only executes user commands but also enables personalized dialogue that is sensitive to the user's emotions. This invention enriches the user experience and enables more natural interaction.
[0514] The following describes the processing flow.
[0515] Step 1:
[0516] The user's device accesses the platform and displays the account registration screen. The user enters their name, email address, and password to create an account. The server receives the input data, checks for duplicates with existing data, and then saves it to the database.
[0517] Step 2:
[0518] The user accesses the login screen using their device and enters their registered email address and password. The server compares this information with the database, and if authentication is successful, it displays the home screen to the user.
[0519] Step 3:
[0520] The server provides the user with a list of AI agent templates on the home screen. The user's device displays this list, and the user can select a template that suits their preferences. The selected template information is then sent to the server.
[0521] Step 4:
[0522] Once a template is selected, the user's device provides visual and audio customization menus. Users can adjust colors, voice tone, and other attributes using sliders and pull-down menus to create their own personalized settings. The server analyzes and stores this customization data in real time.
[0523] Step 5:
[0524] Once customization is complete, the server generates a unique AI agent based on the collected data. The generated agent is then sent from the server to the user's device, where the user can use it.
[0525] Step 6:
[0526] An emotion engine is embedded in the user's device. This emotion engine captures and analyzes the user's voice and facial expressions through the microphone and camera. The server receives this analysis and estimates the user's emotions.
[0527] Step 7:
[0528] Based on the estimated emotional information, the server adjusts the AI agent's response. For example, if it senses that the user is tired, the agent is programmed to send an encouraging message in a cheerful voice. The terminal then presents this response to the user.
[0529] Step 8:
[0530] When a user decides to purchase new visuals or audio through the marketplace, the device displays a purchase screen. The server integrates with an external payment system to complete the purchase process. Once the purchase is confirmed, a download link is provided to the device, allowing the user to access the new content.
[0531] Step 9:
[0532] The server begins the process of transferring existing usage data to the new customized agent. This data migration ensures that users can enjoy a consistent experience with the new agent. The terminal notifies the user when this transfer is complete.
[0533] (Example 2)
[0534] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0535] Traditional AI agent systems have had problems with their inability to flexibly respond to changes in user emotions. Furthermore, their customization interfaces are cumbersome, and changes are not reflected in real time. Additionally, the acquisition of additional content and payment procedures are cumbersome for users.
[0536] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0537] In this invention, the server includes means for providing multiple styles that the user can select, means for the user to modify visual and auditory elements based on the style, and means for performing emotion analysis and adjusting the function's response based on the user's emotional state. This allows the user to easily customize a flexible AI agent that is attuned to their emotions and seamlessly acquire additional content and complete payment procedures.
[0538] "Style" refers to a template that allows users to select the basic structure and characteristics of an AI agent.
[0539] "Visual elements" refer to the components related to the appearance, colors, and design of an AI agent.
[0540] "Acoustic elements" refer to the components related to the quality, tone, and intonation of the AI agent's voice.
[0541] "Functionality" refers to the capabilities and behaviors of an AI agent that have been modified and generated by the user.
[0542] The "marketplace" refers to a virtual trading platform where users can acquire additional content for their AI agents.
[0543] "Device" refers to a device used by a user to interact with an AI agent.
[0544] "Usage information" refers to the data and history accumulated when a user uses the AI agent.
[0545] "Emotion analysis" refers to a technology that infers a user's emotional state from their voice and facial expressions.
[0546] This invention is implemented using a terminal used by the user and a server that processes data. First, the user accesses the platform from the terminal and logs in with their account. The server verifies the user's authentication information and executes the login process.
[0547] Subsequently, the user can select the most suitable AI agent style from several options on the platform. Based on the selected style, an interface is provided on the device to make detailed modifications to the visual and auditory elements. The server receives this modification information and uses the generated AI model to build an AI agent tailored to the user's customizations.
[0548] The AI agent has emotion analysis capabilities, and the device collects voice and posture data during user interaction. This data is analyzed by a server, and the agent's response is adjusted based on the user's emotional state. This technology utilizes natural language processing and image recognition algorithms. For example, if the user has a sad expression, the agent will be configured to use more friendly and encouraging words.
[0549] Users can also acquire additional visual and auditory elements from the marketplace via their devices. The server can transfer purchased materials to the user's device and reflect the new materials without affecting existing agents. Furthermore, a data transfer function ensures that user usage information is carried over to the new agent, maintaining a consistent user experience.
[0550] As a concrete example, consider a scenario where a user selects "Professional Assistant" and sets a formal voice and simple design. In this case, if emotion analysis reveals that the user is experiencing stress, the agent will adjust to provide advice to help them relax.
[0551] An example of a prompt message for a generative AI model would be: "Based on the style of the AI agent selected by the user, analyze the emotions in real time and generate an appropriate response."
[0552] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0553] Step 1:
[0554] Users use their devices to access the platform, enter their account information, and log in.
[0555] Upon receiving the entered authentication information, the server accesses the database to verify its validity. If the information is correct, the user is redirected to the home screen. Specifically, if the user forgets their password, a password reset option is provided.
[0556] Step 2:
[0557] The user selects from several AI agent styles on the home screen.
[0558] The format selected by the terminal is sent to the server. The server returns information about that format to the user and provides data to display a visual preview of the options.
[0559] In terms of specific operation, the user is prompted to decide on the optimal format while viewing a preview.
[0560] Step 3:
[0561] Based on the selected style, users can make detailed modifications to visual and auditory elements on their device. This involves adjusting colors and sound tones using sliders and selectors.
[0562] Correction input is sent from the terminal to the server. The server uses a generation AI model to process the correction data and performs calculations to generate the corrected AI agent.
[0563] The generated data is sent back to the user's terminal, and the corrected results are previewed.
[0564] Step 4:
[0565] The user begins interacting with the generated AI agent.
[0566] The device uses sensors to record the user's voice and facial expressions, and then sends that data to a server.
[0567] The server performs data analysis, uses an emotion analysis engine to identify the user's emotional state, and processes the data so that a generative AI model can generate an appropriate response.
[0568] In terms of specific actions, if the user appears sad, the agent will respond in a gentle tone, saying something like, "Please let me know if there's anything I can help you with."
[0569] Step 5:
[0570] Users can use the marketplace to purchase additional visual and audio elements and download them to their devices.
[0571] Purchase operations are sent from the terminal to the server, and payment is completed by linking with an external payment service.
[0572] The server receives purchase information and processes the data to transfer the necessary content to the user's device.
[0573] The device instantly deploys downloaded content and reflects it in the AI agent. Specifically, the user can check their purchase history and be presented with an option to try new features immediately.
[0574] (Application Example 2)
[0575] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0576] In recent years, digital agents have become increasingly widespread, but many of these agents are unable to empathize with users' emotions during interactions, remaining limited to simply providing information or executing commands. Therefore, there is a need for systems that can provide more natural and personalized communication, especially for the elderly and users who require emotional support. Furthermore, the customizability of agents and the diversity of their content are limited, failing to adequately address the diverse preferences and needs of users.
[0577] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0578] In this invention, the server includes means for providing a user-selectable set of templates, means for the user to customize visual and auditory elements based on the templates, and means for generating and providing a customized agent to the user. This makes it possible to analyze the user's emotional state and adjust the agent's response based on the analysis results. This allows for personalized communication that is attentive to the user's emotions, thereby improving user satisfaction. Furthermore, by utilizing the market to purchase and make downloadable diverse content, a system with scalability to meet diverse user needs is realized.
[0579] A "template" is a set of settings that forms the basis for users to select the appearance and behavior of an agent.
[0580] "Visual" refers to elements related to the appearance and graphics of the user's digital agent.
[0581] "Acoustics" refers to elements related to the voices and sound effects that digital agents emit to users.
[0582] An "agent" is a digital interface that interacts with users, providing information and taking actions based on instructions.
[0583] A "marketplace" is a platform where users can purchase additional content, and where content is provided to expand the user's agent.
[0584] A "server" is a computing system that generates and personalizes digital agents in response to user requests and provides the results to the user.
[0585] "Emotional state" refers to the emotions a user is currently experiencing and serves as foundational information for adjusting the agent's response.
[0586] To implement this invention, a system is required for generating and operating digital agents for users. The server provides the user with multiple selectable templates, from which the user can customize the agent's visual and auditory properties. The server receives the user's selection in real time and generates the customized agent. In doing so, a database is used to consider the user's past usage data, ensuring a consistent user experience.
[0587] The terminal is expected to be a wearable device such as smart glasses. On the terminal, the agent interacts with the user. During this interaction, the terminal's camera and microphone are used to collect the user's facial expressions and voice data. This data is sent to a server, where an emotion engine analyzes the user's emotional state. Based on the analysis results, the agent adjusts its voice and tone of voice, enabling it to respond in a way that is empathetic to the user's stress and anxiety. For implementing the emotion engine, it is recommended to use services such as Google Cloud's Vision API or Amazon Transcribe.
[0588] As a concrete example, imagine an elderly person exchanging daily information with an agent through smart glasses, and the displayed agent gently speaking to the user according to their emotions and offering stress-reducing advice. In this case, the agent might gently encourage the user by saying, "I recommend you relax today. Let's take a few deep breaths."
[0589] An example of a prompt message would be, "Based on the elderly person's facial expressions and voice data, continue the conversation in a relaxed tone and provide advice to help them have a comfortable day." This allows for the use of a generative AI model that adjusts the agent's response.
[0590] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0591] Step 1:
[0592] The user accesses the server using a terminal and selects their preferred digital agent from several templates. The user ID and template ID are provided as input, which the server receives. Based on the template, visual and auditory options are generated and presented to the user. As output, the visual and auditory customization options are displayed on the user's terminal.
[0593] Step 2:
[0594] The user customizes the visual and auditory settings on their device and sends this information to the server. The server processes this customization data in real time as input. The final customization data is saved as output, forming the basis for generating the customized agent. A generative AI model is then used to generate prompts that create an agent based on the customizations.
[0595] Step 3:
[0596] The server utilizes a generative AI model to generate an agent based on stored customized data and delivers it to the user's terminal. Inputs include customized data and generated prompt messages. Output is a digital agent with matching visual and auditory elements displayed on the user's terminal.
[0597] Step 4:
[0598] The device continuously captures the user's facial expressions and voice using its camera and microphone, and sends this data to the server. The input includes voice and image data, which the server analyzes using an emotion engine. The output is an estimated result of the user's emotional state, and this data is used to adjust the agent's response. Specifically, data processing and emotion analysis are performed using Google Cloud's Vision API and Amazon Transcribe.
[0599] Step 5:
[0600] The server adjusts the agent's voice and tone of voice based on the inferred emotional state to tailor its response to the user. The emotion analysis results are included as input, and the adjusted voice response from the agent is played back on the user's device as output. Specifically, the response will be in a relaxed tone to reduce the user's anxiety and stress.
[0601] Step 6:
[0602] The system records the agent's responses to the user and saves them as data to help improve future interactions. Inputs include the agent's responses and user feedback. Outputs include accumulated reference data to improve the consistency of responses and the personalized experience in subsequent interactions.
[0603] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0604] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0605] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0606] [Fourth Embodiment]
[0607] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0608] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0609] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0610] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0611] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0612] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0613] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0614] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0615] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0616] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0617] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0618] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0619] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0620] This invention provides a system that allows users to customize AI agents according to their preferences and enjoy a personalized experience. Users can select an agent from a variety of templates and customize its appearance and voice down to the smallest detail. The system's operation is described below.
[0621] First, the user accesses the platform using their device and creates an account. During this process, they enter the necessary authentication information, and upon successful login, they are redirected to the home screen.
[0622] Next, the server provides the user with several AI agent templates. The user browses through them and selects their preferred template. The selected template displays visual and voice settings menus, which the user uses to customize it.
[0623] Once customization is complete, the server collects the configuration information and generates a customized AI agent for that specific user. This generated agent is then provided to the user's device and ready for interaction.
[0624] Furthermore, by utilizing the marketplace, users can purchase additional visuals and audio created by other creators. The server manages this content and makes the corresponding data available for download to the user's device once the user's purchase is confirmed.
[0625] Purchased agents and customized data are transferred to new agents using the user's past usage data, ensuring a consistent user experience for continued use.
[0626] For example, a user can select a "female character with a pop music style," adjust the pitch and speed of her voice, and change the color of her appearance to their preferred color. This agent is then generated with the selected custom voice and appearance and runs on the user's device.
[0627] In this way, the present invention enables users to obtain an experience tailored to their individual needs through an AI agent. Furthermore, it allows for deeper levels of personalization by leveraging unique customization options.
[0628] The following describes the processing flow.
[0629] Step 1:
[0630] The user's device accesses the platform and displays an account creation screen. Here, the user enters the required information, such as name, email address, and password, to create an account. The server receives the entered data, verifies its accuracy, and then stores it in the database.
[0631] Step 2:
[0632] Authentication is performed by accessing the login screen using the user's device and entering the registered email address and password. The server compares the entered information with the database information, and if successful, redirects the user to the home screen.
[0633] Step 3:
[0634] The server provides the user with a list of multiple AI agent templates on the home screen. The user's device receives this list and displays a template selection screen. The user can then choose their preferred template.
[0635] Step 4:
[0636] Once a template is selected, the user's device displays a visual and audio customization menu. The user can customize the details using tools such as sliders and color selectors. The server records the user's changes in real time and saves the configuration data.
[0637] Step 5:
[0638] Once the user has finished customizing their settings, the server generates a unique AI agent based on that configuration data. This generated agent is sent from the server to the user's device, which the user can then download and begin interacting with.
[0639] Step 6:
[0640] Users are directed to the marketplace where they can browse additional visual and audio content. Once a user decides to purchase, the server interacts with an external payment processing service to complete the transaction. After the purchase is confirmed, a download link is provided, making the content available on the user's device.
[0641] Step 7:
[0642] The server analyzes past agent usage data and transfers it to a new, customized agent. The user's device receives this data and verifies that consistent settings and conversation history are maintained on the new agent.
[0643] (Example 1)
[0644] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0645] Traditional intelligent agent customization had the problem of not being able to adequately address the individual needs of users. Furthermore, the complexity of the operation when making various custom settings made it difficult for users to enjoy intuitive and flexible customization. In addition, existing systems did not allow for smooth retrieval of additional information or transfer of history, resulting in an inconsistent user experience.
[0646] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0647] In this invention, the server includes means for providing visual and auditory information as multiple information sources selectable by the user, means for individually configuring the selected information sources via the user's terminal, and means for the server to generate an intelligent agent corresponding to the user's individual configuration information. This allows the user to intuitively select information sources from multiple options, easily configure them individually, and enjoy an intelligent agent tailored to their individual needs.
[0648] "Information base" refers to a diverse set of visual and auditory features that users can select, and it is the foundation that determines the appearance and voice of an intelligent agent.
[0649] An "intelligent agent" is an interactive program generated by the server based on the user's individual settings, providing functions tailored to the user's specific needs.
[0650] A "trading marketplace" is an online platform where users can purchase additional information and extensions, offering them a variety of options.
[0651] "History transfer" is the process of applying past information created or used by the user to a new intelligent agent, ensuring a consistent user experience.
[0652] "Financial transactions through external services" refers to payment and post-processing activities conducted through services provided by third parties in connection with purchases on trading markets.
[0653] This invention provides a system that allows users to customize intelligent agents according to their preferences and enjoy a personalized experience. Users access the system using a network-connected terminal and utilize the services through that terminal.
[0654] The server retrieves multiple visual and auditory information sources from the database and provides them to the user. The user can browse these information sources through the terminal interface and select the desired ones. For the selected information sources, the user uses the terminal to configure individual settings and set the visual and auditory details.
[0655] Examples of devices used include personal computers, tablets, and smartphones. The server uses a generative AI model to generate an intelligent agent based on the user's settings. This process integrates selected visual and auditory information and performs data editing and analysis within the program.
[0656] Furthermore, users can obtain additional information and enhancements through the trading market provided by the server. The content selected by the user can be downloaded to the device and applied to the intelligent agent. During this process, the server references the user's past usage history and transfers necessary historical information to the new intelligent agent to provide a consistent experience.
[0657] As a concrete example, a user can select a "female character with a pop music style," adjust the pitch and speed of her voice, and change the color of her appearance to their preferred color. This agent is then generated with the selected custom voice and appearance and runs on the user's device.
[0658] Examples of prompt messages include: "Provide specific customization options based on the template selected by the user," and "Describe the process by which the server collects customization settings and generates a new AI agent."
[0659] This system allows users to intuitively customize and enjoy a personalized intelligent agent tailored to their specific needs.
[0660] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0661] Step 1:
[0662] The user accesses the system using their device, creates an account, and logs in. As input, the user provides an email address and password. The server then verifies the authentication information and identifies the user. As output, the home screen is displayed on the user's device.
[0663] Step 2:
[0664] The server provides the user with multiple visual and auditory information sources. The user browses the information sources via an interface on the terminal and selects the desired information sources. The selected information sources are sent to the server as input. Based on this, the server processes the selected information sources and returns them to the user. The information sources selected by the user are displayed on the terminal as output.
[0665] Step 3:
[0666] The user customizes the selected information base. Specifically, they adjust the character's appearance and voice using the settings screen on their device. The input includes the customization parameters specified by the user. The server receives these parameters, analyzes the data using a generative AI model, and generates an intelligent agent. As output, a preview of the customized agent is displayed on the user's device.
[0667] Step 4:
[0668] The server formally generates an intelligent agent based on the customized information and provides it to the user's terminal. The user's final customized data is used as input. The server performs data calculations and sends the generation results to the user's terminal. As output, the fully generated intelligent agent becomes available for use on the terminal.
[0669] Step 5:
[0670] Users can obtain additional information through the trading market. Here, the server provides market information upon user request. Input includes the user's trading requests and payment information. The server processes this information and, if approved, provides purchase data. Output is the additional information downloaded to the user's terminal and reflected in the intelligent agent.
[0671] Step 6:
[0672] The server performs a procedure to inherit past agent usage data. The server receives the user's historical data as input. The server analyzes this data and applies it to the latest agent. The output is an intelligent agent running on the user's terminal, guaranteeing a consistent user experience.
[0673] (Application Example 1)
[0674] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0675] In customizing information processing agents, it is necessary for users to have a personalized experience down to the smallest detail according to their preferences, and to provide a friendly and attractive machine by adjusting the robot's personality and appearance. To achieve this goal, an easily accessible electronic marketplace and a smooth purchasing process are required.
[0676] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0677] In this invention, the server includes means for providing a user with a set of templates to choose from, means for the user to adjust visual and auditory features based on the templates, and means for generating and providing the adjusted information processing agent to the user. This makes it possible to easily realize the customized experience desired by the user and provide a superior user experience.
[0678] A "user" is an individual or legal entity that uses this system to customize information processing agents and enjoy a personalized experience.
[0679] A "template" is a design template that includes pre-defined visual and auditory features, serving as the basis for users to customize information processing agents.
[0680] "Visual features" refer to attributes related to the appearance of an information processing agent, including elements such as color and design.
[0681] "Speech characteristics" refer to the attributes of the voice emitted by an information processing agent, including tone, speed, and intonation.
[0682] An "information processing agent" is a program or machine device that has user-customized visual and auditory characteristics and is designed to perform a specific task.
[0683] An "electronic marketplace" is an online platform where users can purchase additional information and data, and it provides content to expand the customization of agents.
[0684] "Usage data" refers to the history and configuration information of when an existing information processing agent was used, and this information is carried over to the new agent.
[0685] "Individuality" refers to the uniqueness and expressiveness of a robot, and its characteristics are adjusted in appearance and behavior according to the user's preferences.
[0686] In an embodiment of this invention, a user first accesses the system using a dedicated interface and creates their own account. After creating an account, the user can select an information processing agent from a variety of templates. The server provides these templates and allows customization of the selected visual and auditory features. Customization involves setting auditory features using a speech synthesis library (e.g., pyttsx3) and displaying visual features using a graphics library (e.g., pygame).
[0687] The server generates a customized agent and delivers it to the user's terminal. The platform on the terminal provides the ability to purchase additional information data through an electronic marketplace, and the purchased information is designed to be transferred to the terminal immediately. This allows the user to give the customized robot a friendly voice and personality.
[0688] As a concrete example, the AI agent of a cleaning robot can be personalized with a Hawaiian theme and act as a friendly character. Users can select specific voices, such as, "Aloha! Have a great day, I'll start by cleaning the kitchen!" Such customization makes using the robot more enjoyable.
[0689] An example of a prompt to input into the generating AI model is, "Explain how to personalize the AI agent for a cleaning robot with a Hawaiian theme and create a user-friendly character."
[0690] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0691] Step 1:
[0692] The user accesses the system using their terminal and creates an account. The user's authentication information (e.g., username, password) is required as input. The server receives this information, stores it in the database, and generates a new user account. A message confirming successful account creation is displayed on the user's terminal as output.
[0693] Step 2:
[0694] The server presents the user with templates for multiple information processing agents. User account information is required as input. The server verifies the user's access permissions and prepares a list of templates. The list of templates is displayed on the user's terminal as output.
[0695] Step 3:
[0696] The user selects a template and customizes its visual and auditory features. The input requires the user's selected template and customizations (e.g., color scheme, voice tone). The device records this information and displays a realistic preview in real time. The output displays the specific results customized by the user.
[0697] Step 4:
[0698] The server generates a new information processing agent based on the customized settings and makes it available for download to the user's terminal. User customization information is required as input. The server builds the agent and generates compiled data based on this information. As output, a download link is displayed on the user's terminal.
[0699] Step 5:
[0700] The user purchases additional information data from an electronic marketplace. The input requires the user's payment information and selected product information. The server integrates with an external payment service and approves the purchase upon successful payment. The output is a purchase success notification displayed on the user's device, and the product data becomes available for download.
[0701] Step 6:
[0702] The server transfers usage data from the existing agent to the new agent. The server requires usage history data from the existing agent as input. It analyzes this data and synchronizes it with the new agent. The output is a consistent user experience on the new agent, and a notification of the transfer completion is displayed on the user's device.
[0703] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0704] This invention provides a platform that allows users to customize AI agents according to their preferences, and by combining it with an emotion engine, it constructs a system that enables responses based on the user's emotions. The operation of this system is described below.
[0705] First, the user accesses the platform using their device, registers an account, and logs in. The server verifies the user's credentials and guides them to the home screen. Next, the user can select their preferred AI agent template from several options and customize its visuals and voice in detail on their device. The server receives the user's customization information in real time and generates a unique AI agent.
[0706] The generated AI agent is not only provided to the user's device, but also incorporates an emotion engine. During the interaction between the user and the agent, the device senses the user's voice and facial expressions. The emotion engine analyzes these inputs and estimates the user's emotional state. The server adjusts the agent's response according to the estimated emotion. For example, if the user seems sad, the agent will be configured to use a more friendly voice and write encouraging words.
[0707] Furthermore, the system includes a feature that allows users to transfer usage data from existing agents to new user customizations, providing a consistent user experience. Additional content from the marketplace is managed by the server, and users can purchase and download new visuals and sounds to their devices.
[0708] As a concrete example, consider a scenario where a user selects a "business assistant," sets a formal voice, and chooses visuals with somewhat subdued colors. If the emotion engine analyzes that the user is experiencing stress, the agent will be set to offer advice to alleviate stress or to converse in a relaxed tone.
[0709] Such a multi-functional AI agent not only executes user commands but also enables personalized dialogue that is sensitive to the user's emotions. This invention enriches the user experience and enables more natural interaction.
[0710] The following describes the processing flow.
[0711] Step 1:
[0712] The user's device accesses the platform and displays the account registration screen. The user enters their name, email address, and password to create an account. The server receives the input data, checks for duplicates with existing data, and then saves it to the database.
[0713] Step 2:
[0714] The user accesses the login screen using their device and enters their registered email address and password. The server compares this information with the database, and if authentication is successful, it displays the home screen to the user.
[0715] Step 3:
[0716] The server provides the user with a list of AI agent templates on the home screen. The user's device displays this list, and the user can select a template that suits their preferences. The selected template information is then sent to the server.
[0717] Step 4:
[0718] Once a template is selected, the user's device provides visual and audio customization menus. Users can adjust colors, voice tone, and other attributes using sliders and pull-down menus to create their own personalized settings. The server analyzes and stores this customization data in real time.
[0719] Step 5:
[0720] Once customization is complete, the server generates a unique AI agent based on the collected data. The generated agent is then sent from the server to the user's device, where the user can use it.
[0721] Step 6:
[0722] An emotion engine is embedded in the user's device. This emotion engine captures and analyzes the user's voice and facial expressions through the microphone and camera. The server receives this analysis and estimates the user's emotions.
[0723] Step 7:
[0724] Based on the estimated emotional information, the server adjusts the AI agent's response. For example, if it senses that the user is tired, the agent is programmed to send an encouraging message in a cheerful voice. The terminal then presents this response to the user.
[0725] Step 8:
[0726] When a user decides to purchase new visuals or audio through the marketplace, the device displays a purchase screen. The server integrates with an external payment system to complete the purchase process. Once the purchase is confirmed, a download link is provided to the device, allowing the user to access the new content.
[0727] Step 9:
[0728] The server begins the process of transferring existing usage data to the new customized agent. This data migration ensures that users can enjoy a consistent experience with the new agent. The terminal notifies the user when this transfer is complete.
[0729] (Example 2)
[0730] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0731] Traditional AI agent systems have had problems with their inability to flexibly respond to changes in user emotions. Furthermore, their customization interfaces are cumbersome, and changes are not reflected in real time. Additionally, the acquisition of additional content and payment procedures are cumbersome for users.
[0732] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0733] In this invention, the server includes means for providing multiple styles that the user can select, means for the user to modify visual and auditory elements based on the style, and means for performing emotion analysis and adjusting the function's response based on the user's emotional state. This allows the user to easily customize a flexible AI agent that is attuned to their emotions and seamlessly acquire additional content and complete payment procedures.
[0734] "Style" refers to a template that allows users to select the basic structure and characteristics of an AI agent.
[0735] "Visual elements" refer to the components related to the appearance, colors, and design of an AI agent.
[0736] "Acoustic elements" refer to the components related to the quality, tone, and intonation of the AI agent's voice.
[0737] "Functionality" refers to the capabilities and behaviors of an AI agent that have been modified and generated by the user.
[0738] The "marketplace" refers to a virtual trading platform where users can acquire additional content for their AI agents.
[0739] "Device" refers to a device used by a user to interact with an AI agent.
[0740] "Usage information" refers to the data and history accumulated when a user uses the AI agent.
[0741] "Emotion analysis" refers to a technology that infers a user's emotional state from their voice and facial expressions.
[0742] This invention is implemented using a terminal used by the user and a server that processes data. First, the user accesses the platform from the terminal and logs in with their account. The server verifies the user's authentication information and executes the login process.
[0743] Subsequently, the user can select the most suitable AI agent style from several options on the platform. Based on the selected style, an interface is provided on the device to make detailed modifications to the visual and auditory elements. The server receives this modification information and uses the generated AI model to build an AI agent tailored to the user's customizations.
[0744] The AI agent has emotion analysis capabilities, and the device collects voice and posture data during user interaction. This data is analyzed by a server, and the agent's response is adjusted based on the user's emotional state. This technology utilizes natural language processing and image recognition algorithms. For example, if the user has a sad expression, the agent will be configured to use more friendly and encouraging words.
[0745] Users can also acquire additional visual and auditory elements from the marketplace via their devices. The server can transfer purchased materials to the user's device and reflect the new materials without affecting existing agents. Furthermore, a data transfer function ensures that user usage information is carried over to the new agent, maintaining a consistent user experience.
[0746] As a concrete example, consider a scenario where a user selects "Professional Assistant" and sets a formal voice and simple design. In this case, if emotion analysis reveals that the user is experiencing stress, the agent will adjust to provide advice to help them relax.
[0747] An example of a prompt message for a generative AI model would be: "Based on the style of the AI agent selected by the user, analyze the emotions in real time and generate an appropriate response."
[0748] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0749] Step 1:
[0750] Users use their devices to access the platform, enter their account information, and log in.
[0751] Upon receiving the entered authentication information, the server accesses the database to verify its validity. If the information is correct, the user is redirected to the home screen. Specifically, if the user forgets their password, a password reset option is provided.
[0752] Step 2:
[0753] The user selects from several AI agent styles on the home screen.
[0754] The format selected by the terminal is sent to the server. The server returns information about that format to the user and provides data to display a visual preview of the options.
[0755] In terms of specific operation, the user is prompted to decide on the optimal format while viewing a preview.
[0756] Step 3:
[0757] Based on the selected style, users can make detailed modifications to visual and auditory elements on their device. This involves adjusting colors and sound tones using sliders and selectors.
[0758] Correction input is sent from the terminal to the server. The server uses a generation AI model to process the correction data and performs calculations to generate the corrected AI agent.
[0759] The generated data is sent back to the user's terminal, and the corrected results are previewed.
[0760] Step 4:
[0761] The user begins interacting with the generated AI agent.
[0762] The device uses sensors to record the user's voice and facial expressions, and then sends that data to a server.
[0763] The server performs data analysis, uses an emotion analysis engine to identify the user's emotional state, and processes the data so that a generative AI model can generate an appropriate response.
[0764] In terms of specific actions, if the user appears sad, the agent will respond in a gentle tone, saying something like, "Please let me know if there's anything I can help you with."
[0765] Step 5:
[0766] Users can use the marketplace to purchase additional visual and audio elements and download them to their devices.
[0767] Purchase operations are sent from the terminal to the server, and payment is completed by linking with an external payment service.
[0768] The server receives purchase information and processes the data to transfer the necessary content to the user's device.
[0769] The device instantly deploys downloaded content and reflects it in the AI agent. Specifically, the user can check their purchase history and be presented with an option to try new features immediately.
[0770] (Application Example 2)
[0771] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0772] In recent years, digital agents have become increasingly widespread, but many of these agents are unable to empathize with users' emotions during interactions, remaining limited to simply providing information or executing commands. Therefore, there is a need for systems that can provide more natural and personalized communication, especially for the elderly and users who require emotional support. Furthermore, the customizability of agents and the diversity of their content are limited, failing to adequately address the diverse preferences and needs of users.
[0773] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0774] In this invention, the server includes means for providing a user-selectable set of templates, means for the user to customize visual and auditory elements based on the templates, and means for generating and providing a customized agent to the user. This makes it possible to analyze the user's emotional state and adjust the agent's response based on the analysis results. This allows for personalized communication that is attentive to the user's emotions, thereby improving user satisfaction. Furthermore, by utilizing the market to purchase and make downloadable diverse content, a system with scalability to meet diverse user needs is realized.
[0775] A "template" is a set of settings that forms the basis for users to select the appearance and behavior of an agent.
[0776] "Visual" refers to elements related to the appearance and graphics of the user's digital agent.
[0777] "Acoustics" refers to elements related to the voices and sound effects that digital agents emit to users.
[0778] An "agent" is a digital interface that interacts with users, providing information and taking actions based on instructions.
[0779] A "marketplace" is a platform where users can purchase additional content, and where content is provided to expand the user's agent.
[0780] A "server" is a computing system that generates and personalizes digital agents in response to user requests and provides the results to the user.
[0781] "Emotional state" refers to the emotions a user is currently experiencing and serves as foundational information for adjusting the agent's response.
[0782] To implement this invention, a system is required for generating and operating digital agents for users. The server provides the user with multiple selectable templates, from which the user can customize the agent's visual and auditory properties. The server receives the user's selection in real time and generates the customized agent. In doing so, a database is used to consider the user's past usage data, ensuring a consistent user experience.
[0783] The terminal is expected to be a wearable device such as smart glasses. On the terminal, the agent interacts with the user. During this interaction, the terminal's camera and microphone are used to collect the user's facial expressions and voice data. This data is sent to a server, where an emotion engine analyzes the user's emotional state. Based on the analysis results, the agent adjusts its voice and tone of voice, enabling it to respond in a way that is empathetic to the user's stress and anxiety. For implementing the emotion engine, it is recommended to use services such as Google Cloud's Vision API or Amazon Transcribe.
[0784] As a concrete example, imagine an elderly person exchanging daily information with an agent through smart glasses, and the displayed agent gently speaking to the user according to their emotions and offering stress-reducing advice. In this case, the agent might gently encourage the user by saying, "I recommend you relax today. Let's take a few deep breaths."
[0785] An example of a prompt message would be, "Based on the elderly person's facial expressions and voice data, continue the conversation in a relaxed tone and provide advice to help them have a comfortable day." This allows for the use of a generative AI model that adjusts the agent's response.
[0786] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0787] Step 1:
[0788] The user accesses the server using a terminal and selects their preferred digital agent from several templates. The user ID and template ID are provided as input, which the server receives. Based on the template, visual and auditory options are generated and presented to the user. As output, the visual and auditory customization options are displayed on the user's terminal.
[0789] Step 2:
[0790] The user customizes the visual and auditory settings on their device and sends this information to the server. The server processes this customization data in real time as input. The final customization data is saved as output, forming the basis for generating the customized agent. A generative AI model is then used to generate prompts that create an agent based on the customizations.
[0791] Step 3:
[0792] The server utilizes a generative AI model to generate an agent based on stored customized data and delivers it to the user's terminal. Inputs include customized data and generated prompt messages. Output is a digital agent with matching visual and auditory elements displayed on the user's terminal.
[0793] Step 4:
[0794] The device continuously captures the user's facial expressions and voice using its camera and microphone, and sends this data to the server. The input includes voice and image data, which the server analyzes using an emotion engine. The output is an estimated result of the user's emotional state, and this data is used to adjust the agent's response. Specifically, data processing and emotion analysis are performed using Google Cloud's Vision API and Amazon Transcribe.
[0795] Step 5:
[0796] The server adjusts the agent's voice and tone of voice based on the inferred emotional state to tailor its response to the user. The emotion analysis results are included as input, and the adjusted voice response from the agent is played back on the user's device as output. Specifically, the response will be in a relaxed tone to reduce the user's anxiety and stress.
[0797] Step 6:
[0798] The system records the agent's responses to the user and saves them as data to help improve future interactions. Inputs include the agent's responses and user feedback. Outputs include accumulated reference data to improve the consistency of responses and the personalized experience in subsequent interactions.
[0799] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0800] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0801] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0802] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0803] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0804] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0805] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0806] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0807] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0808] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0809] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0810] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0811] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0812] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0813] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0814] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0815] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0816] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0817] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0818] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0819] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0820] The following is further disclosed regarding the embodiments described above.
[0821] (Claim 1)
[0822] A means of providing users with multiple templates to choose from,
[0823] A means for users to customize visuals and audio based on templates,
[0824] A means of generating and providing customized agents to users,
[0825] A means of providing a marketplace that allows users to purchase additional content,
[0826] A means to make purchased content downloadable to the user's device,
[0827] A means of transferring usage data from an existing agent to a new agent,
[0828] A system that includes this.
[0829] (Claim 2)
[0830] The system according to claim 1, further comprising means for saving user changes in real time during the customization process.
[0831] (Claim 3)
[0832] The system according to claim 1, further comprising means for linking with an external payment service in the user's purchase procedure.
[0833] "Example 1"
[0834] (Claim 1)
[0835] A means for providing visual and auditory information as multiple information sources that the user can select,
[0836] A means for individually configuring the information base selected via the user's terminal,
[0837] A means by which the server generates an intelligent agent corresponding to the user's individual configuration information,
[0838] A means of providing the generated intelligent agent to the user's terminal,
[0839] A means of providing users with the ability to obtain and utilize additional information through the trading market,
[0840] A means to make the purchased additional information adaptable to the user's device,
[0841] A means of transferring the history of previously used information to a new intelligent agent,
[0842] A system that includes this.
[0843] (Claim 2)
[0844] The system according to claim 1, further comprising means for immediately reflecting and saving user modifications during the customization process.
[0845] (Claim 3)
[0846] The system according to claim 1, further comprising means for conducting financial transactions through external services in the user acquisition process.
[0847] "Application Example 1"
[0848] (Claim 1)
[0849] A means of providing users with multiple templates to choose from,
[0850] A means by which the user adjusts visual and auditory features based on a template,
[0851] A means for generating and providing a customized information processing agent to the user,
[0852] A means of providing an electronic marketplace that enables users to purchase additional information data,
[0853] A means to enable the transfer of purchased information data to the user's device,
[0854] A method for transferring usage data from an existing agent to a new agent,
[0855] Means to customize the robot's personality and appearance,
[0856] A system that includes this.
[0857] (Claim 2)
[0858] The system according to claim 1, further comprising means for immediately saving user changes during the customization process.
[0859] (Claim 3)
[0860] The system according to claim 1, further comprising means for linking with external payment methods in the user's purchase procedure.
[0861] "Example 2 of combining an emotion engine"
[0862] (Claim 1)
[0863] A means of providing the user with multiple selectable formats,
[0864] Means for users to modify visual and auditory elements based on a style,
[0865] Means for building and providing the corrected functionality to users,
[0866] A means of providing a market that enables users to acquire additional elements,
[0867] Means for making the acquired elements transferable to the user's device,
[0868] A means of transferring usage information from existing functions to new functions,
[0869] A means for performing emotion analysis and adjusting the function's response based on the user's emotional state,
[0870] A system that includes this.
[0871] (Claim 2)
[0872] The system according to claim 1, further comprising means for immediately saving user changes during the customization process.
[0873] (Claim 3)
[0874] The system according to claim 1, further comprising means for linking with external payment technology in the user acquisition procedure.
[0875] "Application example 2 when combining with an emotional engine"
[0876] (Claim 1)
[0877] A means of providing users with multiple templates to choose from,
[0878] A means for users to customize visuals and sounds based on templates,
[0879] A means of generating and providing customized agents to users,
[0880] A means of providing a marketplace that enables users to purchase additional content,
[0881] A means of making purchased content downloadable to the user's device,
[0882] A means of transferring usage data from an existing agent to a new agent,
[0883] A means for analyzing the user's emotional state and adjusting the agent's response based on the analysis results,
[0884] A system that includes this.
[0885] (Claim 2)
[0886] The system according to claim 1, further comprising means for saving user changes in real time during the customization process.
[0887] (Claim 3)
[0888] The system according to claim 1, further comprising means for linking with an external payment service in the user's purchase procedure. [Explanation of Symbols]
[0889] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of providing users with multiple templates to choose from, A means by which the user adjusts visual and auditory features based on a template, A means for generating and providing a customized information processing agent to the user, A means of providing an electronic marketplace that enables users to purchase additional information data, A means to enable the transfer of purchased information data to the user's device, A method for transferring usage data from an existing agent to a new agent, Means to customize the robot's personality and appearance, A system that includes this.
2. The system according to claim 1, further comprising means for immediately saving user changes during the customization process.
3. The system according to claim 1, further comprising means for linking with external payment methods in the user's purchase procedure.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A