system
The system optimizes meeting schedules using AI and provides real-time audio-to-text summaries, addressing the challenge of meeting overload, reducing stress, and enhancing work efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Employees' schedules are filled with meetings, making it difficult to secure lunch breaks, leading to high stress and decreased work efficiency, with challenges in determining important meetings and managing time effectively.
A system that collects user schedule information, optimizes meeting times using AI, generates an avatar for virtual participation, and converts audio to text for real-time summaries, allowing efficient meeting engagement and reduced workload.
Enables employees to concentrate on important meetings, reduce stress, and improve work efficiency by managing time effectively and grasping meeting content quickly.
Smart Images

Figure 2026070203000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, due to the spread of web meetings, especially for management positions, employees' schedules are filled with meetings, and it is difficult to secure a lunch break. As a result, employees are under high stress and work efficiency has decreased. In addition, the existence of many meetings makes it difficult to determine which meetings to attend, and as a result, a lot of time may be spent on meetings with low importance. To solve this problem, it is necessary to provide an environment in which employees can participate in meetings efficiently and with reduced burden.
Means for Solving the Problems
[0005] This invention provides a means to collect user schedule information using an information processing device, optimize meeting times using an artificial intelligence model, and propose meeting times that meet the user's preferences. Furthermore, it generates an avatar that acts as the user's representative in a virtual space to support real-time participation in meetings. It also provides a means for users to efficiently grasp the content of meetings by converting audio information during meetings into text data in real time, analyzing it, and creating summaries. As a result, employees can concentrate on important meetings, reducing stress and improving work efficiency.
[0006] An "information processing device" is a computer device used to manage user schedule information and collect and process necessary data.
[0007] "Schedule information" refers to data containing dates and events that users use for schedule management.
[0008] An "artificial intelligence model" is a program structure that learns based on specific input data and makes predictions or suggestions to optimize the user's meeting time.
[0009] A "terminal" is an electronic device equipped with output and input devices for direct user operation.
[0010] A "virtual space" is a simulation environment created within a computer system that can be visually manipulated by the user.
[0011] An "avatar" refers to a character or image that acts as a proxy for a user in a virtual space.
[0012] "Speech-to-text conversion" is a technical process that converts audio information from a meeting into text data.
[0013] A "summary" is a shortened form of information created by extracting the most important parts from multiple pieces of information or data. [Brief explanation of the drawing]
[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention provides a system that enables users to participate in meetings efficiently and reduce their workload. Specific embodiments for carrying out this invention are described below.
[0036] The server manages user schedule information by retrieving user-registered schedule data from a database. Based on the retrieved schedule information, it uses an artificial intelligence model to calculate the optimal time allocation, taking into account the user's desired meeting and break times. The server sends this calculation result to the terminal, and the schedule is finalized after the user confirms and approves it.
[0037] The user's device displays schedule suggestions received from the server as notifications and sends approvals or feedback to the server based on the user's input. Approved schedules are reflected in the user's calendar application.
[0038] For virtual conference participation, the server provides an avatar generation function, and the user's chosen avatar participates in the conference. The terminal detects the user's voice and text input in real time and reflects it in the avatar. In the virtual conference, the avatar speaks on behalf of the user and visualizes the conference content as needed.
[0039] Furthermore, the audio from the meeting is instantly converted into text data by the server and stored in a database. Based on this text data, the server analyzes the meeting content, extracts summaries and key topics, and provides them to the user via their terminal. This allows the user to quickly grasp the main points of the meeting and use them to their advantage in subsequent work.
[0040] As a concrete example, imagine a manager user who wants to adjust their schedule to ensure they have a lunch break. The AI then adjusts the schedule, allowing them to hold an afternoon meeting. In this scenario, the user participates in the meeting using an avatar, and after the meeting, they can quickly make decisions about subsequent processes based on a summary provided by the server.
[0041] By implementing this invention, users can reduce the burden of attending meetings, manage their time efficiently, and ultimately expect an improvement in overall work productivity.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] The server retrieves schedule information from the database using the user ID. This data includes the user's schedule and meeting information.
[0045] Step 2:
[0046] The server inputs the acquired schedule information into an artificial intelligence model and generates suggestions to optimize meeting and break times. The user's desired conditions (e.g., ensuring a lunch break) are taken into consideration.
[0047] Step 3:
[0048] The server sends the generated schedule proposal to the terminal and notifies the user. The terminal receives this information and displays it as a pop-up on the screen.
[0049] Step 4:
[0050] The user operates the terminal to review the proposed schedule and enter their approval or any necessary modifications. This input is sent to the server and used to finalize or further adjust the schedule.
[0051] Step 5:
[0052] The server records approved schedules in a database and synchronizes them with calendar applications, etc.
[0053] Step 6:
[0054] Once the meeting begins, the server generates and places the user's selected avatar within the virtual space. The avatar then prepares to reflect the user's statements.
[0055] Step 7:
[0056] The terminal receives user voice or text input in real time and sends it to the server via the network. The server then reflects this in the avatar's movements to facilitate communication within the meeting.
[0057] Step 8:
[0058] During the meeting, the server converts audio data into text, providing real-time transcription. This text is stored in a database and can be accessed even after the meeting ends.
[0059] Step 9:
[0060] Once the meeting ends, the server analyzes the text data and generates a summary. It extracts important topics and statements and sends them to the terminal in a visually easy-to-understand format.
[0061] Step 10:
[0062] Users can view summary information provided on their devices to review meeting content. Based on this, they can then adjust their work and make decisions.
[0063] (Example 1)
[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0065] Traditional scheduling management systems had problems with efficient time allocation, as users had to individually set and adjust meeting times, which was time-consuming. Furthermore, in virtual meeting environments, incomplete information transmission due to a lack of real-world communication was a challenge. Additionally, the lack of features to effectively summarize and notify meeting information made it difficult to grasp the key points of the meeting.
[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0067] In this invention, the server includes means for acquiring user schedule information from a data storage device, means for calculating an optimal time allocation according to the user's wishes using a generation AI model based on the acquired schedule information, means for notifying the user terminal of the calculated time allocation and receiving approval or correction feedback from the user, means for reflecting the approved time allocation in a calendar application program, means for generating a selected display image in a virtual space and having it operate as a user's proxy, means for reflecting voice and text information transmitted from the user via the terminal in the display image in real time, means for converting voice information collected during a meeting into text information in real time using voice processing means, information processing means for analyzing the text information and summarizing the meeting content, and means for providing the summarized information to the user through notification means. As a result, the user can efficiently manage their schedule, communicate smoothly in a virtual environment, and quickly grasp the key points of a meeting.
[0068] A "user" refers to an individual or group that has registered to use an information system.
[0069] "Schedule information" refers to a series of date, time, and event-related data registered and managed by the user.
[0070] A "data storage device" refers to a storage medium or device that can store and retrieve information.
[0071] A "generative AI model" is an artificial intelligence algorithm used to generate the optimal solution by performing complex calculations and reasoning based on human instructions.
[0072] "Time allocation" is a concept that refers to the optimal distribution of time according to specific activities or requirements.
[0073] A "terminal" is an electronic device used by a user to input and retrieve information, and usually refers to a computer or smartphone.
[0074] A "calendar application program" is a software application used for managing dates, times, and events.
[0075] A "virtual space" is a digital environment created by computer technology that users can interact with.
[0076] "Displayed image" refers to a visible substitute that acts as a representative of the user within a virtual space, such as an avatar.
[0077] "Speech processing means" refers to technologies and devices that analyze speech data and perform necessary conversions and analyses.
[0078] "Textual information" refers to information in text format converted from audio or other data.
[0079] "Information processing means" refers to technologies and devices that have the ability to analyze data and process it into a useful form according to a specific purpose.
[0080] "Notification means" refers to technologies and devices for informing a user of specific information visually or audibly.
[0081] This invention is a system that allows users to efficiently manage their schedules and effectively participate in meetings in a virtual environment. The following describes a specific implementation of this system.
[0082] The server first retrieves the user's schedule information from its data storage device. A common database management system can be used for this process. For example, the database stores the user's schedule information, and the server retrieves the necessary data using SQL queries.
[0083] Next, the server uses a generative AI model based on the acquired schedule information to calculate the optimal time allocation that meets the user's preferences. Various artificial intelligence platforms can be used as generative AI models in this process. In this process, the AI is instructed with a prompt such as, "Please tell me how to get the optimal meeting schedule suggested to the user," and the calculation results are obtained.
[0084] The user's device receives notifications of optimal time allocation sent from the server and displays them to the user. The device, using a smartphone or PC, provides an interface for the user to check, approve, or modify their schedule. This feedback is resent to the server, and the finally approved schedule is reflected in the calendar application program.
[0085] In virtual conference participation, the server generates a display image based on the user's selection. This display image acts as a proxy for the user in the virtual space and is implemented using common computer graphics technology. The user's terminal sends voice and text to the server via the terminal device, and the server reflects this in the display image in real time.
[0086] Furthermore, audio data from meetings is instantly converted into text information by a server using audio processing equipment. This text information is then analyzed by information processing equipment to extract summaries and key topics. The results of this analysis are provided to the user through notification systems, allowing them to quickly grasp the main points of the meeting. This entire system enables users to efficiently manage their schedules and effectively utilize the content of meetings.
[0087] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0088] Step 1:
[0089] The server retrieves the user's schedule information from the database. An SQL query is issued using the user's ID as input, and the corresponding schedule information is output from the database. Specifically, the server executes the query "SELECT FROM schedules WHERE user_id = User ID" and stores the retrieved data in its internal storage.
[0090] Step 2:
[0091] The server uses an AI model to calculate the optimal time allocation based on the acquired schedule information. The input is the schedule information and the prompt "Please suggest the optimal time allocation desired by the user." The output is the time allocation suggestion generated by the AI. The AI model performs internal calculations and generates a timetable based on the parameters to be terminated.
[0092] Step 3:
[0093] The server sends the generated time allocation to the user terminal. The input is the time allocation obtained in step 2, and the output is the notification to the user terminal. Specifically, the server uses a REST API to send an HTTP POST request to the user terminal, and the terminal receives the data and displays the notification on the screen.
[0094] Step 4:
[0095] The terminal displays proposals from the server to the user and accepts approval or modification from the user. The input is the time allocation proposal received from the server, and the output is the user's feedback. Specifically, the terminal screen displays "Approve" or "Modify" buttons, and after the user makes a selection, the feedback is sent to the server.
[0096] Step 5:
[0097] The server reflects the approved schedule in the calendar application program. The input is the user's approval feedback, and the output is the registration of the schedule in the calendar application program. Specifically, the server uses the calendar API and sends a request to add the schedule after authenticating with OAuth 2.0.
[0098] Step 6:
[0099] The server generates a display image of the user's selection for participation in a virtual meeting. The input is the user's selection information, and the output is the display image in the virtual space. The server uses a virtual environment engine to generate a display image based on the user's selection and reflect it in the virtual space.
[0100] Step 7:
[0101] The terminal sends voice and text from the user to the server, which is then reflected in the display in real time. Input is voice data and text information, and output is real-time reflection in a virtual space. The terminal collects information from the microphone and keyboard and sends it to the server sequentially.
[0102] Step 8:
[0103] The server converts audio information from a meeting into text, analyzes it, and generates a summary. The input is audio data, and the output is the summarized meeting content. Specifically, the server uses an audio processing engine to convert the audio into text, analyzes the text to extract key points, and generates a summary.
[0104] (Application Example 1)
[0105] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0106] In today's business environment, activities in virtual spaces and interactions using digital people are becoming increasingly important. However, existing technologies have not adequately addressed the challenges of user schedule management and the efficiency of actual conversations. This invention aims to provide a system that allows users to efficiently engage in activities in virtual spaces, reduce their workload, and conduct effective conversations.
[0107] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0108] In this invention, the server includes means for collecting user activity information managed by an information processing device, means for proposing a time allocation according to the user's wishes using an artificial intelligence model generated to optimize work time based on the activity information, and means for generating a digital person that acts as a proxy for the user in a virtual space. This enables the user to streamline activities in the virtual environment, reduce the user's workload, and improve the quality of interaction.
[0109] An "information processing device" is a computer device used to manage user activity information and schedules, and to control the entire system.
[0110] "Activity information" refers to data about various actions and schedules that users perform on a daily basis.
[0111] An "artificial intelligence model" is an AI algorithm used to analyze data and make optimal suggestions in order to optimize user activities.
[0112] "Time allocation" refers to the distribution of time to efficiently carry out activities and tasks according to the user's preferences.
[0113] A "virtual space" is a digital environment created by computers that is different from the real world.
[0114] A "digital character" is an avatar that acts on behalf of the user to perform various interactions in a virtual space.
[0115] "Real-time" refers to processing and responses occurring almost instantaneously, meaning that information is provided to the user without delay.
[0116] "Text data" refers to data obtained by converting spoken or oral information into text format.
[0117] A "summary" is a short, concise compilation of information or data, extracted from a large amount of information or data.
[0118] "Interaction" refers to the interaction or communication between a user and a system, or between users themselves.
[0119] This invention is a system that enables users to efficiently manage their activities and make the most of interactions in virtual space. Its main components include an information processing device, an artificial intelligence model, a digital person generation function, and a speech-to-text conversion function.
[0120] The server acts as an information processing device, collecting and managing user activity information. Considering user preferences and activity patterns, it optimizes work time using a generative AI model. To achieve this, it analyzes data using machine learning libraries such as TENSORFLOW® and presents suggestions in a user-readable format. This allows users to create a more rational schedule.
[0121] The server also generates a digital persona to represent the user in the virtual space. This digital persona operates using real-time communication technologies such as WebRTC, reflecting the user's voice and text instructions. This allows users to communicate in a virtual environment without having to participate physically.
[0122] Furthermore, the server converts audio information into text data in real time. By using the Google® Cloud Speech-to-Text API, the acquired audio information is efficiently converted into text and a summary of key topics is generated. This allows users to quickly grasp the main points of a conversation and use that information to aid in subsequent decision-making.
[0123] As a concrete example, suppose user A has a virtual meeting scheduled for the morning. The server proposes a pre-optimized schedule, and once user A confirms it, the digital persona joins the meeting and assists with the discussion. After the meeting ends, the server automatically summarizes the key points and provides them to user A. It is also possible to improve the quality of the digital persona's responses by using a prompt such as, "Suggest how the avatar should respond appropriately to the customer."
[0124] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0125] Step 1:
[0126] The server collects user activity information from a database. The input consists of schedules and daily activity data registered by the user. This collected information serves as the basis for optimization algorithms that are then processed. This prepares the system for more efficient management of user activity.
[0127] Step 2:
[0128] The server uses an artificial intelligence model to optimize work time based on the collected activity information. The collected activity information is used as input, and this information is analyzed to propose the optimal time allocation for each activity. This data processing is a necessary step to propose an efficient schedule to the user.
[0129] Step 3:
[0130] The server notifies the user's terminal of the optimized time allocation. The input is the optimized schedule generated by the artificial intelligence model. As output, the optimized schedule is displayed on the user's terminal, and the user can review its contents and submit approval or revision suggestions. This finalizes a feasible schedule that reflects the user's feedback.
[0131] Step 4:
[0132] The server generates a digital persona to represent the user in a virtual space. Inputs consist of approved schedule information and user instructions. The generated digital persona receives voice and text instructions from the user in real time and operates within the virtual environment. This allows the user to continue activities within the virtual environment without direct participation.
[0133] Step 5:
[0134] The server converts speech from digital characters and user voice input into text data in real time. The input is voice data from interactions in the virtual space. The output is the generated text data. This process allows the voice information to be saved in text format, making it easy to refer to later.
[0135] Step 6:
[0136] The server analyzes the text data and summarizes the key points. Text data is used as input, and summarized text is generated as output. This process allows users to concisely understand vast amounts of information and make important decisions quickly.
[0137] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0138] This invention provides a system that supports users in efficiently participating in meetings while also considering their emotions during communication. Specific embodiments for carrying out this invention are described below.
[0139] The server retrieves information from a database to manage user schedules and uses an artificial intelligence model to optimize meeting times. This process takes into account the user's preferences. The proposed schedule is then notified to the terminal by the server and can be reviewed or modified by the user. Once the user approves the schedule, the server stores it in the database and synchronizes it with the relevant applications.
[0140] For virtual meeting participation, the server generates an avatar, and the terminal reflects the user's voice and text input in real time. Furthermore, by incorporating an emotion engine, the server detects emotions from the user's facial expressions and voice patterns and reflects them in the avatar's expressions and movements. This allows the avatar to behave more naturally and human-like.
[0141] Audio data from meetings is converted into text data in real time by a server. The server analyzes the meeting content based on this text data and creates a summary. Furthermore, it analyzes the mood of the meeting using emotional information detected by an emotion engine and reports the impression and atmosphere to the user in a visualized form.
[0142] As a concrete example, consider a scenario where an employee participates in an online meeting. The server adjusts the user's schedule, ensuring a lunch break before setting up the afternoon meeting. During the meeting, an emotion engine determines whether the employee is relaxed or stressed, and appropriately changes the avatar's facial expressions accordingly. After the meeting, along with a summary of the meeting content, the server points out the periods when positive emotions were most prevalent, suggesting areas for improvement and successes for the user.
[0143] By implementing this invention, users can expect to maintain efficient and emotionally engaging communication, leading to improved work performance and results.
[0144] The following describes the processing flow.
[0145] Step 1:
[0146] The server retrieves schedule information from the database using the user ID. This information includes the user's schedule and desired break times.
[0147] Step 2:
[0148] The server uses an artificial intelligence model to calculate the optimal meeting time based on the acquired schedule information. In doing so, it takes into account the user's preferences.
[0149] Step 3:
[0150] The server sends the calculated schedule proposal to the terminal and notifies the user. The terminal displays this information and provides the user with options.
[0151] Step 4:
[0152] Users use their devices to review the proposed schedule and approve or modify it. User input is fed back to the server.
[0153] Step 5:
[0154] The server finalizes the schedule based on user feedback and saves it to the database. It also synchronizes with the calendar app as needed.
[0155] Step 6:
[0156] When a meeting begins, the server generates and places an avatar in the virtual space. The avatar acts as a representative of the user and prepares to receive the user's messages.
[0157] Step 7:
[0158] The device collects the user's voice and text input in real time and sends it to a server via the network. The server reflects the received content in the avatar and takes appropriate actions during the meeting.
[0159] Step 8:
[0160] The server uses an emotion engine to analyze the user's voice patterns and facial expressions, acquiring emotional information in real time. This information is reflected in the avatar's movements, naturally expressing the user's emotions.
[0161] Step 9:
[0162] During the meeting, the server converts audio information into text data in real time and saves it to the database. The converted data can be retrieved even after the meeting has ended.
[0163] Step 10:
[0164] After the meeting ends, the server analyzes the text data and sentiment information to generate a summary of the meeting content and sentiment. The generated summary and analysis report are provided to the user via their terminal.
[0165] Step 11:
[0166] Users can view the provided summary information and sentiment analysis reports to use for reviewing their work and preparing for future meetings.
[0167] (Example 2)
[0168] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0169] In today's busy environment, optimizing schedules and communicating with consideration for emotions are crucial for users to participate in meetings efficiently and smoothly. However, conventional schedule management systems and video conferencing tools have struggled to adequately reflect users' preferences and emotions when proposing schedules and communicating. Furthermore, it has been difficult to grasp the content of meetings and obtain meaningful feedback through emotional analysis during meetings. This invention aims to solve these problems and improve users' work efficiency and the quality of communication.
[0170] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0171] In this invention, the server includes means for acquiring schedule information by a device that collects and manages user schedule information, means for optimizing and proposing meeting times that reflect the user's desired conditions using a generated artificial intelligence model, and means for notifying the user's information device of the proposed meeting times and receiving approval or modification instructions. This enables the user to efficiently manage their schedule and participate in meetings optimally. Furthermore, by reflecting voice and text data onto a virtual entity representing the user and operating it in real time, it realizes natural communication that takes emotions into consideration. In addition, by converting voice data during meetings into text data in real time and providing a summary of the meeting content and the results of sentiment analysis during the meeting, the user can review their impression of the meeting and areas for improvement, and receive feedback that leads to improved work performance.
[0172] A "user" refers to an individual or member of an organization that uses the system, and is the entity that participates in meetings and manages schedules.
[0173] "Schedule information" refers to information related to the user's schedule, including data on the time and conditions of meetings, events, etc.
[0174] "Information equipment" refers to electronic devices used by users, including smartphones and personal computers.
[0175] A "data management device" is a device that provides means for storing, managing, and synchronizing data, and includes cloud servers and databases.
[0176] "Existence" refers to a digital character or avatar that acts on behalf of the user within a virtual realm.
[0177] The "generated artificial intelligence model" is a system composed of algorithms used for data analysis and prediction, and is particularly responsible for optimizing meeting times by taking user conditions into consideration.
[0178] "Real-time" refers to the characteristic that data processing and communication occur almost instantly, resulting in a state where the system responds to input without delay.
[0179] "Sentiment analysis" refers to the process of inferring and analyzing a user's emotional state based on their tone of voice and the content of their writing.
[0180] "Voice data conversion" refers to the process of converting voice information into text information, and is performed using speech recognition technology.
[0181] A "mood report" refers to information that visualizes the emotional state of participants and the atmosphere of a meeting, and is provided to users as feedback.
[0182] This invention provides a system that enables users to efficiently participate in meetings using information devices and to achieve emotionally sensitive communication. Specifically, a server collects user schedule information and manages it with a data management device. This allows the server to have a detailed understanding of the user's schedule information and use a generated artificial intelligence model to optimize meeting times based on the user's preferences.
[0183] The server notifies the user's information device of the optimized schedule and receives approval or modification instructions from the user. The modified schedule is stored by the server on a data management device and synchronized with the relevant application software. This allows the user to participate in meetings in a planned manner.
[0184] Furthermore, to support participation in virtual meetings, the server generates an avatar that acts as a proxy for the user. This avatar reflects the voice and text data transmitted from the user via the terminal and operates in real time. In addition, by incorporating an emotion analysis engine, the server detects the user's emotional state from their voice tone and speech content and immediately reflects it in the avatar's facial expressions. This function enables smooth communication that takes emotions into consideration.
[0185] During the meeting, the server converts audio data into text data in real time. This process uses speech recognition software. Based on the text data, the server summarizes the meeting content and analyzes the emotional information acquired during the meeting. The analysis results are provided to the user as a mood report, allowing the user to review the meeting's impression and key points and receive feedback.
[0186] For example, if a user has multiple meetings scheduled in the morning, the server uses a generative AI model to adjust morning and afternoon meeting times and suggest ways to ensure a lunch break. An example of a prompt message given to the generative AI model would be, "Optimize the schedule by adjusting afternoon meetings based on the user's preferences." This allows users to manage their schedules with less stress and achieve better work efficiency.
[0187] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0188] Step 1:
[0189] The server retrieves user schedule information from the data management device. The inputs used are the user's ID and associated past schedule information. Based on the retrieved schedule information, a data analysis algorithm extracts user trends and preferences, which are then provided to the AI model. The output generates the initial dataset necessary for schedule optimization.
[0190] Step 2:
[0191] The server uses the generated artificial intelligence model to optimize the schedule. In this process, the user's conditions and optimization goals are input to the generated AI model as prompts. The AI analyzes the input data and calculates the optimal time for meetings. As output, candidate meeting times are listed. As a concrete example, the prompt "Please schedule the meeting in the afternoon, avoiding lunchtime" is used.
[0192] Step 3:
[0193] The server notifies the user's information device of an optimized meeting schedule. The input is a list of meeting times suggested by a generative AI model. The server formats this as notification data and sends it to the user's terminal. As output, the suggested meeting schedule is displayed on the user's screen. The user reviews the presented dates and times and makes corrections as needed.
[0194] Step 4:
[0195] The user reviews the schedule and enters approval or modification information. The terminal receives the user's input and sends the modified schedule data back to the server. The input is the user's approval or modification instruction, and the output is the updated schedule data. This allows the user to manage schedules that meet their own criteria.
[0196] Step 5:
[0197] In preparation for a meeting in a virtual space, the server generates a user avatar. The inputs for this process are the user profile and past meeting data. The server configures the avatar's basic behavior and appearance, and outputs the data as a 3D model. During the meeting, the terminal reflects the user's voice and text onto the avatar, enabling real-time operation.
[0198] Step 6:
[0199] The server converts audio data into text data in real time. The input is an audio stream from a meeting, and the server uses speech recognition software to convert it into text. The output is the converted text data, which is used later for summarization and sentiment analysis. Specifically, it transcribes user speech into text, helping to visualize the meeting content.
[0200] Step 7:
[0201] The server summarizes meeting content and performs sentiment analysis based on text data. Input consists of real-time text data and elements for sentiment analysis. A data analysis algorithm processes this data to generate summarized meeting content and sentiment data. Output includes a visualized mood report and meeting summaries. Users can use this to assess the effectiveness of the meeting and improve future meetings.
[0202] (Application Example 2)
[0203] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0204] In today's commercial environment, online and offline portals are becoming increasingly integrated, but the user experience in virtual spaces is often limited. This results in a lack of deep interaction that considers user emotions, making it challenging to provide a service comparable to the experience in a physical store. Furthermore, there is a demand for flexible scheduling and personalized service that takes user emotions into account during meetings and business negotiations, but this is difficult to achieve with current technology.
[0205] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0206] In this invention, the server includes information processing means for collecting user schedule information, data analysis means for proposing a time that suits the user's preferences using an algorithm for optimizing meeting times based on the schedule information, and emotion analysis means for detecting the user's emotions and reflecting them in a corresponding digital representation. This enables a personalized virtual space experience that takes the user's emotions into consideration.
[0207] A "communication system" is a general term for technologies and devices used to send and receive data between information processing devices and terminals located in different places.
[0208] "Information processing means" refers to hardware or software functions for collecting, processing, and converting user data and information into a useful form.
[0209] A "data analysis tool" is a program or processing method that analyzes collected data and derives useful information according to a specific purpose.
[0210] "Emotional analysis means" refers to a technology or process that determines a user's psychological state and analyzes their emotions based on their voice, facial expressions, etc.
[0211] A "digital character" is a virtual entity that acts on behalf of the user on a computer and interacts with them.
[0212] "Generation means" refers to the process or device used to create new digital content based on specified data.
[0213] "Data conversion means" refers to a process or device that converts data in one format into another format, making it usable.
[0214] "Information display means" refers to technology or devices for visually providing data, results, or information to a user through an interface.
[0215] This invention is a system that enables effective communication in a virtual space while taking into account the user's emotions. The server collects the user's schedule information using information processing means and proposes the optimal meeting time using an algorithm based on the schedule information. To reflect the user's emotions, emotion analysis means analyzes the user's emotions from their voice and facial expressions and generates a corresponding digital representation. This digital representation is reflected in a digital character that acts as a proxy for the user in the virtual space, supporting interaction with the user.
[0216] In this process, the server acquires user emotion data in real time using Azure® Kinect or an equivalent emotion detection device. Based on this data, a digital character is generated that provides voice guidance using speech synthesis services such as Amazon Polly, enabling natural communication with the user.
[0217] For example, if a user wants advice on a specific product in a virtual store, the system can determine from the user's cautious expression that they are hesitant about purchasing it. A digital character then gently encourages and supports the user by saying, "This product has been very well-received by many users. Why don't you give it a try?"
[0218] An example of a prompt is, "How can an AI model create a shopping guide that matches the user's emotions?" This prompt provides guidance to the generative AI model for generating appropriate guides and content.
[0219] In this way, the invention enables communication that takes emotions into account, leading to a more fulfilling experience for users.
[0220] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0221] Step 1:
[0222] The server collects schedule information from the user's smart glasses. As input, it retrieves data from the user's calendar or schedule. The server analyzes this data and prepares to suggest optimal meeting times based on the user's schedule patterns and preferences. The analyzed schedule information is then generated as output.
[0223] Step 2:
[0224] The server uses emotion analysis to acquire facial and voice data in real time from sensors installed in the user's smart glasses. The user's facial and voice data are provided as input. The server analyzes this data using an emotion detection algorithm to determine the user's emotional state. The user's emotional state is then generated as output.
[0225] Step 3:
[0226] The server determines the actions of the digital character in the virtual space based on the emotional state. The emotional state obtained in the previous step is used as input. This allows the digital character to reflect expressions and actions that correspond to the user's emotions, supporting interaction with the user. Action instructions for the digital character are generated as output.
[0227] Step 4:
[0228] The device controls a digital character and displays actions corresponding to the user's smart glasses. Input is provided as instructions for the digital character's actions from a server. The device uses a speech synthesis service such as Amazon Polly, allowing the digital character to make suggestions and explanations to the user in a natural voice. Output is generated as audio and visual feedback to the user.
[0229] Step 5:
[0230] The user reviews suggestions from the server and interacts with digital characters, then decides on actions such as making a purchase or attending a meeting. Input comes from information provided by the digital characters. Based on the user's actions and choices, the terminal records this information and sends it back to the server for further data analysis. Output is generated, which includes the user's selection information.
[0231] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0232] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0233] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0234] [Second Embodiment]
[0235] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0236] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0237] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0238] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0239] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0240] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0241] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0242] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0243] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0244] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0245] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0246] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0247] This invention provides a system that enables users to participate in meetings efficiently and reduce their workload. Specific embodiments for carrying out this invention are described below.
[0248] The server manages user schedule information by retrieving user-registered schedule data from a database. Based on the retrieved schedule information, it uses an artificial intelligence model to calculate the optimal time allocation, taking into account the user's desired meeting and break times. The server sends this calculation result to the terminal, and the schedule is finalized after the user confirms and approves it.
[0249] The user's device displays schedule suggestions received from the server as notifications and sends approvals or feedback to the server based on the user's input. Approved schedules are reflected in the user's calendar application.
[0250] For virtual conference participation, the server provides an avatar generation function, and the user's chosen avatar participates in the conference. The terminal detects the user's voice and text input in real time and reflects it in the avatar. In the virtual conference, the avatar speaks on behalf of the user and visualizes the conference content as needed.
[0251] Furthermore, the audio from the meeting is instantly converted into text data by the server and stored in a database. Based on this text data, the server analyzes the meeting content, extracts summaries and key topics, and provides them to the user via their terminal. This allows the user to quickly grasp the main points of the meeting and use them to their advantage in subsequent work.
[0252] As a concrete example, imagine a manager user who wants to adjust their schedule to ensure they have a lunch break. The AI then adjusts the schedule, allowing them to hold an afternoon meeting. In this scenario, the user participates in the meeting using an avatar, and after the meeting, they can quickly make decisions about subsequent processes based on a summary provided by the server.
[0253] By implementing this invention, users can reduce the burden of attending meetings, manage their time efficiently, and ultimately expect an improvement in overall work productivity.
[0254] The following describes the processing flow.
[0255] Step 1:
[0256] The server retrieves schedule information from the database using the user ID. This data includes the user's schedule and meeting information.
[0257] Step 2:
[0258] The server inputs the acquired schedule information into an artificial intelligence model and generates suggestions to optimize meeting and break times. The user's desired conditions (e.g., ensuring a lunch break) are taken into consideration.
[0259] Step 3:
[0260] The server sends the generated schedule proposal to the terminal and notifies the user. The terminal receives this information and displays it as a pop-up on the screen.
[0261] Step 4:
[0262] The user operates the terminal to review the proposed schedule and enter their approval or any necessary modifications. This input is sent to the server and used to finalize or further adjust the schedule.
[0263] Step 5:
[0264] The server records approved schedules in a database and synchronizes them with calendar applications, etc.
[0265] Step 6:
[0266] Once the meeting begins, the server generates and places the user's selected avatar within the virtual space. The avatar then prepares to reflect the user's statements.
[0267] Step 7:
[0268] The terminal receives user voice or text input in real time and sends it to the server via the network. The server then reflects this in the avatar's movements to facilitate communication within the meeting.
[0269] Step 8:
[0270] During the meeting, the server converts audio data into text, providing real-time transcription. This text is stored in a database and can be accessed even after the meeting ends.
[0271] Step 9:
[0272] Once the meeting ends, the server analyzes the text data and generates a summary. It extracts important topics and statements and sends them to the terminal in a visually easy-to-understand format.
[0273] Step 10:
[0274] Users can view summary information provided on their devices to review meeting content. Based on this, they can then adjust their work and make decisions.
[0275] (Example 1)
[0276] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0277] Traditional scheduling management systems had problems with efficient time allocation, as users had to individually set and adjust meeting times, which was time-consuming. Furthermore, in virtual meeting environments, incomplete information transmission due to a lack of real-world communication was a challenge. Additionally, the lack of features to effectively summarize and notify meeting information made it difficult to grasp the key points of the meeting.
[0278] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0279] In this invention, the server includes means for obtaining the user's schedule information from a data storage device, means for calculating an optimal time allocation according to the user's wishes by utilizing a generation AI model based on the obtained schedule information, means for notifying the calculated time allocation to the user terminal and receiving approval or correction feedback from the user, means for reflecting the approved time allocation in a calendar application program, means for generating a selected display image in a virtual space and operating as a user agent, means for real-time reflecting the voice and text information transmitted from the user via the terminal in the display image, means for converting the voice information collected during a meeting into character information in real time by voice processing means, information processing means for analyzing the character information and summarizing the meeting content, and means for providing the summarized information to the user through a notification means. Thereby, the user can perform efficient schedule management, smooth communication in a virtual environment, and quickly grasp the key points of a meeting.
[0280] The "user" refers to an individual or group registered for using an information system.
[0281] The "schedule information" is a series of data related to dates and times and events registered and managed by the user.
[0282] The "data storage device" refers to a storage medium or device capable of storing and obtaining information.
[0283] The "generation AI model" is an artificial intelligence algorithm used to perform complex calculations and inferences based on human instructions and generate an optimal solution.
[0284] The "time allocation" refers to the concept of optimally allocating time according to specific activities or requirements.
[0285] The "terminal" is an electronic device used by the user to input and obtain information, usually referring to a computer or a smartphone.
[0286] A "calendar application program" is a software application for managing dates, times, and events.
[0287] A "virtual space" is a digital environment generated by computer technology in which users can interact.
[0288] A "display avatar" is a visible substitute that operates on behalf of the user within the virtual space and refers to an avatar or the like.
[0289] "Voice processing means" refers to technologies and devices that analyze voice data and perform necessary conversions and analyses.
[0290] "Character information" refers to information in text format converted from voice or other data.
[0291] "Information processing means" refers to technologies and devices that have the ability to analyze data and process it into a useful form according to specific purposes.
[0292] "Notification means" refers to technologies and devices for visually or auditorily notifying users of specific information.
[0293] This invention is a system that enables users to efficiently manage schedules and effectively participate in meetings in a virtual environment. The following shows specific forms for implementing this system.
[0294] First, the server obtains the user's schedule information from the data storage device. In this process, a general database management system can be used. For example, the user's schedule information is stored in the database, and the server uses SQL queries to obtain the necessary data.
[0295] Next, the server uses a generative AI model based on the acquired schedule information to calculate the optimal time allocation that meets the user's preferences. Various artificial intelligence platforms can be used as generative AI models in this process. In this process, the AI is instructed with a prompt such as, "Please tell me how to get the optimal meeting schedule suggested to the user," and the calculation results are obtained.
[0296] The user's device receives notifications of optimal time allocation sent from the server and displays them to the user. The device, using a smartphone or PC, provides an interface for the user to check, approve, or modify their schedule. This feedback is resent to the server, and the finally approved schedule is reflected in the calendar application program.
[0297] In virtual conference participation, the server generates a display image based on the user's selection. This display image acts as a proxy for the user in the virtual space and is implemented using common computer graphics technology. The user's terminal sends voice and text to the server via the terminal device, and the server reflects this in the display image in real time.
[0298] Furthermore, audio data from meetings is instantly converted into text information by a server using audio processing equipment. This text information is then analyzed by information processing equipment to extract summaries and key topics. The results of this analysis are provided to the user through notification systems, allowing them to quickly grasp the main points of the meeting. This entire system enables users to efficiently manage their schedules and effectively utilize the content of meetings.
[0299] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0300] Step 1:
[0301] The server retrieves the user's schedule information from the database. Using the user ID as input, an SQL query is issued, and the corresponding schedule information is output from the database. As a specific operation, the server executes a query "SELECT FROM schedules WHERE user_id = user ID" and holds the retrieved data in the built-in storage.
[0302] Step 2:
[0303] Based on the retrieved schedule information, the server uses a generative AI model to calculate the optimal time allocation. The inputs are the schedule information and the prompt sentence "Please propose the optimal time allocation desired by the user". The output is a proposal for the time allocation generated by the AI. The AI model performs internal calculations and generates a schedule based on the termination parameters.
[0304] Step 3:
[0305] The server sends the generated time allocation to the user terminal. The input is the time allocation obtained in Step 2, and the output is a notification to the user terminal. As a specific operation, the server uses the REST API to send an HTTP POST request to the user terminal, and the terminal receives the data and displays the notification on the screen.
[0306] Step 4:
[0307] The terminal displays the proposal from the server to the user and accepts approval or modification from the user. The input is the time allocation proposal received from the server, and the output is feedback from the user. The specific operation is to place "Approve" or "Modify" buttons on the terminal screen, receive the user's selection, and send feedback to the server.
[0308] Step 5:
[0309] The server reflects the approved schedule in the calendar application program. The input is the user's approval feedback, and the output is the registration of the schedule in the calendar application program. Specifically, the server uses the calendar API and sends a request to add the schedule after authenticating with OAuth 2.0.
[0310] Step 6:
[0311] The server generates a display image of the user's selection for participation in a virtual meeting. The input is the user's selection information, and the output is the display image in the virtual space. The server uses a virtual environment engine to generate a display image based on the user's selection and reflect it in the virtual space.
[0312] Step 7:
[0313] The terminal sends voice and text from the user to the server, which is then reflected in the display in real time. Input is voice data and text information, and output is real-time reflection in a virtual space. The terminal collects information from the microphone and keyboard and sends it to the server sequentially.
[0314] Step 8:
[0315] The server converts audio information from a meeting into text, analyzes it, and generates a summary. The input is audio data, and the output is the summarized meeting content. Specifically, the server uses an audio processing engine to convert the audio into text, analyzes the text to extract key points, and generates a summary.
[0316] (Application Example 1)
[0317] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0318] In today's business environment, activities in virtual spaces and interactions using digital people are becoming increasingly important. However, existing technologies have not adequately addressed the challenges of user schedule management and the efficiency of actual conversations. This invention aims to provide a system that allows users to efficiently engage in activities in virtual spaces, reduce their workload, and conduct effective conversations.
[0319] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0320] In this invention, the server includes means for collecting user activity information managed by an information processing device, means for proposing a time allocation according to the user's wishes using an artificial intelligence model generated to optimize work time based on the activity information, and means for generating a digital person that acts as a proxy for the user in a virtual space. This enables the user to streamline activities in the virtual environment, reduce the user's workload, and improve the quality of interaction.
[0321] An "information processing device" is a computer device used to manage user activity information and schedules, and to control the entire system.
[0322] "Activity information" refers to data about various actions and schedules that users perform on a daily basis.
[0323] An "artificial intelligence model" is an AI algorithm used to analyze data and make optimal suggestions in order to optimize user activities.
[0324] "Time allocation" refers to the distribution of time to efficiently carry out activities and tasks according to the user's preferences.
[0325] A "virtual space" is a digital environment created by computers that is different from the real world.
[0326] A "digital character" is an avatar that acts on behalf of the user to perform various interactions in a virtual space.
[0327] "Real-time" refers to processing and responses occurring almost instantaneously, meaning that information is provided to the user without delay.
[0328] "Text data" refers to data obtained by converting spoken or oral information into text format.
[0329] A "summary" is a short, concise compilation of information or data, extracted from a large amount of information or data.
[0330] "Interaction" refers to the interaction or communication between a user and a system, or between users themselves.
[0331] This invention is a system that enables users to efficiently manage their activities and make the most of interactions in virtual space. Its main components include an information processing device, an artificial intelligence model, a digital person generation function, and a speech-to-text conversion function.
[0332] The server acts as an information processing device, collecting and managing user activity information. Considering user preferences and activity patterns, it optimizes work time using a generative AI model. To achieve this, it analyzes data using machine learning libraries such as TensorFlow and provides suggestions in a user-readable format. This allows users to create a more rational schedule.
[0333] The server also generates a digital persona to represent the user in the virtual space. This digital persona operates using real-time communication technologies such as WebRTC, reflecting the user's voice and text instructions. This allows users to communicate in a virtual environment without having to participate physically.
[0334] Furthermore, the server converts audio information into text data in real time. By using the Google Cloud Speech-to-Text API, the acquired audio information is efficiently converted into text and a summary of key topics is generated. This allows users to quickly grasp the main points of a conversation and use that information to aid in subsequent decision-making.
[0335] As a concrete example, suppose user A has a virtual meeting scheduled for the morning. The server proposes a pre-optimized schedule, and once user A confirms it, the digital persona joins the meeting and assists with the discussion. After the meeting ends, the server automatically summarizes the key points and provides them to user A. It is also possible to improve the quality of the digital persona's responses by using a prompt such as, "Suggest how the avatar should respond appropriately to the customer."
[0336] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0337] Step 1:
[0338] The server collects user activity information from a database. The input consists of schedules and daily activity data registered by the user. This collected information serves as the basis for optimization algorithms that are then processed. This prepares the system for more efficient management of user activity.
[0339] Step 2:
[0340] The server uses an artificial intelligence model to optimize work time based on the collected activity information. The collected activity information is used as input, and this information is analyzed to propose the optimal time allocation for each activity. This data processing is a necessary step to propose an efficient schedule to the user.
[0341] Step 3:
[0342] The server notifies the user's terminal of the optimized time allocation. The input is the optimized schedule generated by the artificial intelligence model. As output, the optimized schedule is displayed on the user's terminal, and the user can review its contents and submit approval or revision suggestions. This finalizes a feasible schedule that reflects the user's feedback.
[0343] Step 4:
[0344] The server generates a digital persona to represent the user in a virtual space. Inputs consist of approved schedule information and user instructions. The generated digital persona receives voice and text instructions from the user in real time and operates within the virtual environment. This allows the user to continue activities within the virtual environment without direct participation.
[0345] Step 5:
[0346] The server converts speech from digital characters and user voice input into text data in real time. The input is voice data from interactions in the virtual space. The output is the generated text data. This process allows the voice information to be saved in text format, making it easy to refer to later.
[0347] Step 6:
[0348] The server analyzes the text data and summarizes the key points. Text data is used as input, and summarized text is generated as output. This process allows users to concisely understand vast amounts of information and make important decisions quickly.
[0349] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0350] This invention provides a system that supports users in efficiently participating in meetings while also considering their emotions during communication. Specific embodiments for carrying out this invention are described below.
[0351] The server retrieves information from a database to manage user schedules and uses an artificial intelligence model to optimize meeting times. This process takes into account the user's preferences. The proposed schedule is then notified to the terminal by the server and can be reviewed or modified by the user. Once the user approves the schedule, the server stores it in the database and synchronizes it with the relevant applications.
[0352] For virtual meeting participation, the server generates an avatar, and the terminal reflects the user's voice and text input in real time. Furthermore, by incorporating an emotion engine, the server detects emotions from the user's facial expressions and voice patterns and reflects them in the avatar's expressions and movements. This allows the avatar to behave more naturally and human-like.
[0353] Audio data from meetings is converted into text data in real time by a server. The server analyzes the meeting content based on this text data and creates a summary. Furthermore, it analyzes the mood of the meeting using emotional information detected by an emotion engine and reports the impression and atmosphere to the user in a visualized form.
[0354] As a concrete example, consider a scenario where an employee participates in an online meeting. The server adjusts the user's schedule, ensuring a lunch break before setting up the afternoon meeting. During the meeting, an emotion engine determines whether the employee is relaxed or stressed, and appropriately changes the avatar's facial expressions accordingly. After the meeting, along with a summary of the meeting content, the server points out the periods when positive emotions were most prevalent, suggesting areas for improvement and successes for the user.
[0355] By implementing this invention, users can expect to maintain efficient and emotionally engaging communication, leading to improved work performance and results.
[0356] The following describes the processing flow.
[0357] Step 1:
[0358] The server retrieves schedule information from the database using the user ID. This information includes the user's schedule and desired break times.
[0359] Step 2:
[0360] The server uses an artificial intelligence model to calculate the optimal meeting time based on the acquired schedule information. In doing so, it takes into account the user's preferences.
[0361] Step 3:
[0362] The server sends the calculated schedule proposal to the terminal and notifies the user. The terminal displays this information and provides the user with options.
[0363] Step 4:
[0364] Users use their devices to review the proposed schedule and approve or modify it. User input is fed back to the server.
[0365] Step 5:
[0366] The server finalizes the schedule based on user feedback and saves it to the database. It also synchronizes with the calendar app as needed.
[0367] Step 6:
[0368] When a meeting begins, the server generates and places an avatar in the virtual space. The avatar acts as a representative of the user and prepares to receive the user's messages.
[0369] Step 7:
[0370] The device collects the user's voice and text input in real time and sends it to a server via the network. The server reflects the received content in the avatar and takes appropriate actions during the meeting.
[0371] Step 8:
[0372] The server uses an emotion engine to analyze the user's voice patterns and facial expressions, acquiring emotional information in real time. This information is reflected in the avatar's movements, naturally expressing the user's emotions.
[0373] Step 9:
[0374] During the meeting, the server converts audio information into text data in real time and saves it to the database. The converted data can be retrieved even after the meeting has ended.
[0375] Step 10:
[0376] After the meeting ends, the server analyzes the text data and sentiment information to generate a summary of the meeting content and sentiment. The generated summary and analysis report are provided to the user via their terminal.
[0377] Step 11:
[0378] Users can view the provided summary information and sentiment analysis reports to use for reviewing their work and preparing for future meetings.
[0379] (Example 2)
[0380] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0381] In today's busy environment, optimizing schedules and communicating with consideration for emotions are crucial for users to participate in meetings efficiently and smoothly. However, conventional schedule management systems and video conferencing tools have struggled to adequately reflect users' preferences and emotions when proposing schedules and communicating. Furthermore, it has been difficult to grasp the content of meetings and obtain meaningful feedback through emotional analysis during meetings. This invention aims to solve these problems and improve users' work efficiency and the quality of communication.
[0382] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0383] In this invention, the server includes means for acquiring schedule information by a device that collects and manages user schedule information, means for optimizing and proposing meeting times that reflect the user's desired conditions using a generated artificial intelligence model, and means for notifying the user's information device of the proposed meeting times and receiving approval or modification instructions. This enables the user to efficiently manage their schedule and participate in meetings optimally. Furthermore, by reflecting voice and text data onto a virtual entity representing the user and operating it in real time, it realizes natural communication that takes emotions into consideration. In addition, by converting voice data during meetings into text data in real time and providing a summary of the meeting content and the results of sentiment analysis during the meeting, the user can review their impression of the meeting and areas for improvement, and receive feedback that leads to improved work performance.
[0384] A "user" refers to an individual or member of an organization that uses the system, and is the entity that participates in meetings and manages schedules.
[0385] "Schedule information" refers to information related to the user's schedule, including data on the time and conditions of meetings, events, etc.
[0386] "Information equipment" refers to electronic devices used by users, including smartphones and personal computers.
[0387] A "data management device" is a device that provides means for storing, managing, and synchronizing data, and includes cloud servers and databases.
[0388] "Existence" refers to a digital character or avatar that acts on behalf of the user within a virtual realm.
[0389] The "generated artificial intelligence model" is a system composed of algorithms used for data analysis and prediction, and is particularly responsible for optimizing meeting times by taking user conditions into consideration.
[0390] "Real-time" refers to the characteristic that data processing and communication occur almost instantly, resulting in a state where the system responds to input without delay.
[0391] "Sentiment analysis" refers to the process of inferring and analyzing a user's emotional state based on their tone of voice and the content of their writing.
[0392] "Voice data conversion" refers to the process of converting voice information into text information, and is performed using speech recognition technology.
[0393] A "mood report" refers to information that visualizes the emotional state of participants and the atmosphere of a meeting, and is provided to users as feedback.
[0394] This invention provides a system that enables users to efficiently participate in meetings using information devices and to achieve emotionally sensitive communication. Specifically, a server collects user schedule information and manages it with a data management device. This allows the server to have a detailed understanding of the user's schedule information and use a generated artificial intelligence model to optimize meeting times based on the user's preferences.
[0395] The server notifies the user's information device of the optimized schedule and receives approval or modification instructions from the user. The modified schedule is stored by the server on a data management device and synchronized with the relevant application software. This allows the user to participate in meetings in a planned manner.
[0396] Furthermore, to support participation in virtual meetings, the server generates an avatar that acts as a proxy for the user. This avatar reflects the voice and text data transmitted from the user via the terminal and operates in real time. In addition, by incorporating an emotion analysis engine, the server detects the user's emotional state from their voice tone and speech content and immediately reflects it in the avatar's facial expressions. This function enables smooth communication that takes emotions into consideration.
[0397] During the meeting, the server converts audio data into text data in real time. This process uses speech recognition software. Based on the text data, the server summarizes the meeting content and analyzes the emotional information acquired during the meeting. The analysis results are provided to the user as a mood report, allowing the user to review the meeting's impression and key points and receive feedback.
[0398] For example, if a user has multiple meetings scheduled in the morning, the server uses a generative AI model to adjust morning and afternoon meeting times and suggest ways to ensure a lunch break. An example of a prompt message given to the generative AI model would be, "Optimize the schedule by adjusting afternoon meetings based on the user's preferences." This allows users to manage their schedules with less stress and achieve better work efficiency.
[0399] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0400] Step 1:
[0401] The server retrieves user schedule information from the data management device. The inputs used are the user's ID and associated past schedule information. Based on the retrieved schedule information, a data analysis algorithm extracts user trends and preferences, which are then provided to the AI model. The output generates the initial dataset necessary for schedule optimization.
[0402] Step 2:
[0403] The server uses the generated artificial intelligence model to optimize the schedule. In this process, the user's conditions and optimization goals are input to the generated AI model as prompts. The AI analyzes the input data and calculates the optimal time for meetings. As output, candidate meeting times are listed. As a concrete example, the prompt "Please schedule the meeting in the afternoon, avoiding lunchtime" is used.
[0404] Step 3:
[0405] The server notifies the user's information device of an optimized meeting schedule. The input is a list of meeting times suggested by a generative AI model. The server formats this as notification data and sends it to the user's terminal. As output, the suggested meeting schedule is displayed on the user's screen. The user reviews the presented dates and times and makes corrections as needed.
[0406] Step 4:
[0407] The user reviews the schedule and enters approval or modification information. The terminal receives the user's input and sends the modified schedule data back to the server. The input is the user's approval or modification instruction, and the output is the updated schedule data. This allows the user to manage schedules that meet their own criteria.
[0408] Step 5:
[0409] In preparation for a meeting in a virtual space, the server generates a user avatar. The inputs for this process are the user profile and past meeting data. The server configures the avatar's basic behavior and appearance, and outputs the data as a 3D model. During the meeting, the terminal reflects the user's voice and text onto the avatar, enabling real-time operation.
[0410] Step 6:
[0411] The server converts audio data into text data in real time. The input is an audio stream from a meeting, and the server uses speech recognition software to convert it into text. The output is the converted text data, which is used later for summarization and sentiment analysis. Specifically, it transcribes user speech into text, helping to visualize the meeting content.
[0412] Step 7:
[0413] The server summarizes meeting content and performs sentiment analysis based on text data. Input consists of real-time text data and elements for sentiment analysis. A data analysis algorithm processes this data to generate summarized meeting content and sentiment data. Output includes a visualized mood report and meeting summaries. Users can use this to assess the effectiveness of the meeting and improve future meetings.
[0414] (Application Example 2)
[0415] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0416] In today's commercial environment, online and offline portals are becoming increasingly integrated, but the user experience in virtual spaces is often limited. This results in a lack of deep interaction that considers user emotions, making it challenging to provide a service comparable to the experience in a physical store. Furthermore, there is a demand for flexible scheduling and personalized service that takes user emotions into account during meetings and business negotiations, but this is difficult to achieve with current technology.
[0417] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0418] In this invention, the server includes information processing means for collecting user schedule information, data analysis means for proposing a time that suits the user's preferences using an algorithm for optimizing meeting times based on the schedule information, and emotion analysis means for detecting the user's emotions and reflecting them in a corresponding digital representation. This enables a personalized virtual space experience that takes the user's emotions into consideration.
[0419] A "communication system" is a general term for technologies and devices used to send and receive data between information processing devices and terminals located in different places.
[0420] "Information processing means" refers to hardware or software functions for collecting, processing, and converting user data and information into a useful form.
[0421] A "data analysis tool" is a program or processing method that analyzes collected data and derives useful information according to a specific purpose.
[0422] "Emotional analysis means" refers to a technology or process that determines a user's psychological state and analyzes their emotions based on their voice, facial expressions, etc.
[0423] A "digital character" is a virtual entity that acts on behalf of the user on a computer and interacts with them.
[0424] "Generation means" refers to the process or device used to create new digital content based on specified data.
[0425] "Data conversion means" refers to a process or device that converts data in one format into another format, making it usable.
[0426] "Information display means" refers to technology or devices for visually providing data, results, or information to a user through an interface.
[0427] This invention is a system that enables effective communication in a virtual space while taking into account the user's emotions. The server collects the user's schedule information using information processing means and proposes the optimal meeting time using an algorithm based on the schedule information. To reflect the user's emotions, emotion analysis means analyzes the user's emotions from their voice and facial expressions and generates a corresponding digital representation. This digital representation is reflected in a digital character that acts as a proxy for the user in the virtual space, supporting interaction with the user.
[0428] During this process, the server uses Azure Kinect or equivalent emotion detection devices to acquire user emotion data in real time. Based on this data, a digital character is generated that uses speech synthesis services such as Amazon Polly to provide voice guidance and engage in natural communication with the user.
[0429] For example, if a user wants advice on a specific product in a virtual store, the system can determine from the user's cautious expression that they are hesitant about purchasing it. A digital character then gently encourages and supports the user by saying, "This product has been very well-received by many users. Why don't you give it a try?"
[0430] An example of a prompt is, "How can an AI model create a shopping guide that matches the user's emotions?" This prompt provides guidance to the generative AI model for generating appropriate guides and content.
[0431] In this way, the invention enables communication that takes emotions into account, leading to a more fulfilling experience for users.
[0432] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0433] Step 1:
[0434] The server collects schedule information from the user's smart glasses. As input, it retrieves data from the user's calendar or schedule. The server analyzes this data and prepares to suggest optimal meeting times based on the user's schedule patterns and preferences. The analyzed schedule information is then generated as output.
[0435] Step 2:
[0436] The server uses emotion analysis to acquire facial and voice data in real time from sensors installed in the user's smart glasses. The user's facial and voice data are provided as input. The server analyzes this data using an emotion detection algorithm to determine the user's emotional state. The user's emotional state is then generated as output.
[0437] Step 3:
[0438] The server determines the actions of the digital character in the virtual space based on the emotional state. The emotional state obtained in the previous step is used as input. This allows the digital character to reflect expressions and actions that correspond to the user's emotions, supporting interaction with the user. Action instructions for the digital character are generated as output.
[0439] Step 4:
[0440] The device controls a digital character and displays actions corresponding to the user's smart glasses. Input is provided as instructions for the digital character's actions from a server. The device uses a speech synthesis service such as Amazon Polly, allowing the digital character to make suggestions and explanations to the user in a natural voice. Output is generated as audio and visual feedback to the user.
[0441] Step 5:
[0442] The user reviews suggestions from the server and interacts with digital characters, then decides on actions such as making a purchase or attending a meeting. Input comes from information provided by the digital characters. Based on the user's actions and choices, the terminal records this information and sends it back to the server for further data analysis. Output is generated, which includes the user's selection information.
[0443] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0444] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0445] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0446] [Third Embodiment]
[0447] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0448] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0449] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0450] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0451] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0452] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0453] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0454] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0455] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0456] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0457] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0458] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0459] This invention provides a system that enables users to participate in meetings efficiently and reduce their workload. Specific embodiments for carrying out this invention are described below.
[0460] The server manages user schedule information by retrieving user-registered schedule data from a database. Based on the retrieved schedule information, it uses an artificial intelligence model to calculate the optimal time allocation, taking into account the user's desired meeting and break times. The server sends this calculation result to the terminal, and the schedule is finalized after the user confirms and approves it.
[0461] The user's device displays schedule suggestions received from the server as notifications and sends approvals or feedback to the server based on the user's input. Approved schedules are reflected in the user's calendar application.
[0462] For virtual conference participation, the server provides an avatar generation function, and the user's chosen avatar participates in the conference. The terminal detects the user's voice and text input in real time and reflects it in the avatar. In the virtual conference, the avatar speaks on behalf of the user and visualizes the conference content as needed.
[0463] Furthermore, the audio from the meeting is instantly converted into text data by the server and stored in a database. Based on this text data, the server analyzes the meeting content, extracts summaries and key topics, and provides them to the user via their terminal. This allows the user to quickly grasp the main points of the meeting and use them to their advantage in subsequent work.
[0464] As a concrete example, imagine a manager user who wants to adjust their schedule to ensure they have a lunch break. The AI then adjusts the schedule, allowing them to hold an afternoon meeting. In this scenario, the user participates in the meeting using an avatar, and after the meeting, they can quickly make decisions about subsequent processes based on a summary provided by the server.
[0465] By implementing this invention, users can reduce the burden of attending meetings, manage their time efficiently, and ultimately expect an improvement in overall work productivity.
[0466] The following describes the processing flow.
[0467] Step 1:
[0468] The server retrieves schedule information from the database using the user ID. This data includes the user's schedule and meeting information.
[0469] Step 2:
[0470] The server inputs the acquired schedule information into an artificial intelligence model and generates suggestions to optimize meeting and break times. The user's desired conditions (e.g., ensuring a lunch break) are taken into consideration.
[0471] Step 3:
[0472] The server sends the generated schedule proposal to the terminal and notifies the user. The terminal receives this information and displays it as a pop-up on the screen.
[0473] Step 4:
[0474] The user operates the terminal to review the proposed schedule and enter their approval or any necessary modifications. This input is sent to the server and used to finalize or further adjust the schedule.
[0475] Step 5:
[0476] The server records approved schedules in a database and synchronizes them with calendar applications, etc.
[0477] Step 6:
[0478] Once the meeting begins, the server generates and places the user's selected avatar within the virtual space. The avatar then prepares to reflect the user's statements.
[0479] Step 7:
[0480] The terminal receives user voice or text input in real time and sends it to the server via the network. The server then reflects this in the avatar's movements to facilitate communication within the meeting.
[0481] Step 8:
[0482] During the meeting, the server converts audio data into text, providing real-time transcription. This text is stored in a database and can be accessed even after the meeting ends.
[0483] Step 9:
[0484] Once the meeting ends, the server analyzes the text data and generates a summary. It extracts important topics and statements and sends them to the terminal in a visually easy-to-understand format.
[0485] Step 10:
[0486] Users can view summary information provided on their devices to review meeting content. Based on this, they can then adjust their work and make decisions.
[0487] (Example 1)
[0488] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0489] Traditional scheduling management systems had problems with efficient time allocation, as users had to individually set and adjust meeting times, which was time-consuming. Furthermore, in virtual meeting environments, incomplete information transmission due to a lack of real-world communication was a challenge. Additionally, the lack of features to effectively summarize and notify meeting information made it difficult to grasp the key points of the meeting.
[0490] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0491] In this invention, the server includes means for acquiring user schedule information from a data storage device, means for calculating an optimal time allocation according to the user's wishes using a generation AI model based on the acquired schedule information, means for notifying the user terminal of the calculated time allocation and receiving approval or correction feedback from the user, means for reflecting the approved time allocation in a calendar application program, means for generating a selected display image in a virtual space and having it operate as a user's proxy, means for reflecting voice and text information transmitted from the user via the terminal in the display image in real time, means for converting voice information collected during a meeting into text information in real time using voice processing means, information processing means for analyzing the text information and summarizing the meeting content, and means for providing the summarized information to the user through notification means. As a result, the user can efficiently manage their schedule, communicate smoothly in a virtual environment, and quickly grasp the key points of a meeting.
[0492] A "user" refers to an individual or group that has registered to use an information system.
[0493] "Schedule information" refers to a series of date, time, and event-related data registered and managed by the user.
[0494] A "data storage device" refers to a storage medium or device that can store and retrieve information.
[0495] A "generative AI model" is an artificial intelligence algorithm used to generate the optimal solution by performing complex calculations and reasoning based on human instructions.
[0496] "Time allocation" is a concept that refers to the optimal distribution of time according to specific activities or requirements.
[0497] A "terminal" is an electronic device used by a user to input and retrieve information, and usually refers to a computer or smartphone.
[0498] A "calendar application program" is a software application used for managing dates, times, and events.
[0499] A "virtual space" is a digital environment created by computer technology that users can interact with.
[0500] "Displayed image" refers to a visible substitute that acts as a representative of the user within a virtual space, such as an avatar.
[0501] "Speech processing means" refers to technologies and devices that analyze speech data and perform necessary conversions and analyses.
[0502] "Textual information" refers to information in text format converted from audio or other data.
[0503] "Information processing means" refers to technologies and devices that have the ability to analyze data and process it into a useful form according to a specific purpose.
[0504] "Notification means" refers to technologies and devices for informing a user of specific information visually or audibly.
[0505] This invention is a system that allows users to efficiently manage their schedules and effectively participate in meetings in a virtual environment. The following describes a specific implementation of this system.
[0506] The server first retrieves the user's schedule information from its data storage device. A common database management system can be used for this process. For example, the database stores the user's schedule information, and the server retrieves the necessary data using SQL queries.
[0507] Next, the server uses a generative AI model based on the acquired schedule information to calculate the optimal time allocation that meets the user's preferences. Various artificial intelligence platforms can be used as generative AI models in this process. In this process, the AI is instructed with a prompt such as, "Please tell me how to get the optimal meeting schedule suggested to the user," and the calculation results are obtained.
[0508] The user's device receives notifications of optimal time allocation sent from the server and displays them to the user. The device, using a smartphone or PC, provides an interface for the user to check, approve, or modify their schedule. This feedback is resent to the server, and the finally approved schedule is reflected in the calendar application program.
[0509] In virtual conference participation, the server generates a display image based on the user's selection. This display image acts as a proxy for the user in the virtual space and is implemented using common computer graphics technology. The user's terminal sends voice and text to the server via the terminal device, and the server reflects this in the display image in real time.
[0510] Furthermore, audio data from meetings is instantly converted into text information by a server using audio processing equipment. This text information is then analyzed by information processing equipment to extract summaries and key topics. The results of this analysis are provided to the user through notification systems, allowing them to quickly grasp the main points of the meeting. This entire system enables users to efficiently manage their schedules and effectively utilize the content of meetings.
[0511] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0512] Step 1:
[0513] The server retrieves the user's schedule information from the database. An SQL query is issued using the user's ID as input, and the corresponding schedule information is output from the database. Specifically, the server executes the query "SELECT FROM schedules WHERE user_id = User ID" and stores the retrieved data in its internal storage.
[0514] Step 2:
[0515] The server uses an AI model to calculate the optimal time allocation based on the acquired schedule information. The input is the schedule information and the prompt "Please suggest the optimal time allocation desired by the user." The output is the time allocation suggestion generated by the AI. The AI model performs internal calculations and generates a timetable based on the parameters to be terminated.
[0516] Step 3:
[0517] The server sends the generated time allocation to the user terminal. The input is the time allocation obtained in step 2, and the output is the notification to the user terminal. Specifically, the server uses a REST API to send an HTTP POST request to the user terminal, and the terminal receives the data and displays the notification on the screen.
[0518] Step 4:
[0519] The terminal displays proposals from the server to the user and accepts approval or modification from the user. The input is the time allocation proposal received from the server, and the output is the user's feedback. Specifically, the terminal screen displays "Approve" or "Modify" buttons, and after the user makes a selection, the feedback is sent to the server.
[0520] Step 5:
[0521] The server reflects the approved schedule in the calendar application program. The input is the user's approval feedback, and the output is the registration of the schedule in the calendar application program. Specifically, the server uses the calendar API and sends a request to add the schedule after authenticating with OAuth 2.0.
[0522] Step 6:
[0523] The server generates a display image of the user's selection for participation in a virtual meeting. The input is the user's selection information, and the output is the display image in the virtual space. The server uses a virtual environment engine to generate a display image based on the user's selection and reflect it in the virtual space.
[0524] Step 7:
[0525] The terminal sends voice and text from the user to the server, which is then reflected in the display in real time. Input is voice data and text information, and output is real-time reflection in a virtual space. The terminal collects information from the microphone and keyboard and sends it to the server sequentially.
[0526] Step 8:
[0527] The server converts audio information from a meeting into text, analyzes it, and generates a summary. The input is audio data, and the output is the summarized meeting content. Specifically, the server uses an audio processing engine to convert the audio into text, analyzes the text to extract key points, and generates a summary.
[0528] (Application Example 1)
[0529] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0530] In today's business environment, activities in virtual spaces and interactions using digital people are becoming increasingly important. However, existing technologies have not adequately addressed the challenges of user schedule management and the efficiency of actual conversations. This invention aims to provide a system that allows users to efficiently engage in activities in virtual spaces, reduce their workload, and conduct effective conversations.
[0531] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0532] In this invention, the server includes means for collecting user activity information managed by an information processing device, means for proposing a time allocation according to the user's wishes using an artificial intelligence model generated to optimize work time based on the activity information, and means for generating a digital person that acts as a proxy for the user in a virtual space. This enables the user to streamline activities in the virtual environment, reduce the user's workload, and improve the quality of interaction.
[0533] An "information processing device" is a computer device used to manage user activity information and schedules, and to control the entire system.
[0534] "Activity information" refers to data about various actions and schedules that users perform on a daily basis.
[0535] An "artificial intelligence model" is an AI algorithm used to analyze data and make optimal suggestions in order to optimize user activities.
[0536] "Time allocation" refers to the distribution of time to efficiently carry out activities and tasks according to the user's preferences.
[0537] A "virtual space" is a digital environment created by computers that is different from the real world.
[0538] A "digital character" is an avatar that acts on behalf of the user to perform various interactions in a virtual space.
[0539] "Real-time" refers to processing and responses occurring almost instantaneously, meaning that information is provided to the user without delay.
[0540] "Text data" refers to data obtained by converting spoken or oral information into text format.
[0541] A "summary" is a short, concise compilation of information or data, extracted from a large amount of information or data.
[0542] "Interaction" refers to the interaction or communication between a user and a system, or between users themselves.
[0543] This invention is a system that enables users to efficiently manage their activities and make the most of interactions in virtual space. Its main components include an information processing device, an artificial intelligence model, a digital person generation function, and a speech-to-text conversion function.
[0544] The server acts as an information processing device, collecting and managing user activity information. Considering user preferences and activity patterns, it optimizes work time using a generative AI model. To achieve this, it analyzes data using machine learning libraries such as TensorFlow and provides suggestions in a user-readable format. This allows users to create a more rational schedule.
[0545] The server also generates a digital persona to represent the user in the virtual space. This digital persona operates using real-time communication technologies such as WebRTC, reflecting the user's voice and text instructions. This allows users to communicate in a virtual environment without having to participate physically.
[0546] Furthermore, the server converts audio information into text data in real time. By using the Google Cloud Speech-to-Text API, the acquired audio information is efficiently converted into text and a summary of key topics is generated. This allows users to quickly grasp the main points of a conversation and use that information to aid in subsequent decision-making.
[0547] As a concrete example, suppose user A has a virtual meeting scheduled for the morning. The server proposes a pre-optimized schedule, and once user A confirms it, the digital persona joins the meeting and assists with the discussion. After the meeting ends, the server automatically summarizes the key points and provides them to user A. It is also possible to improve the quality of the digital persona's responses by using a prompt such as, "Suggest how the avatar should respond appropriately to the customer."
[0548] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0549] Step 1:
[0550] The server collects user activity information from a database. The input consists of schedules and daily activity data registered by the user. This collected information serves as the basis for optimization algorithms that are then processed. This prepares the system for more efficient management of user activity.
[0551] Step 2:
[0552] The server uses an artificial intelligence model to optimize work time based on the collected activity information. The collected activity information is used as input, and this information is analyzed to propose the optimal time allocation for each activity. This data processing is a necessary step to propose an efficient schedule to the user.
[0553] Step 3:
[0554] The server notifies the user's terminal of the optimized time allocation. The input is the optimized schedule generated by the artificial intelligence model. As output, the optimized schedule is displayed on the user's terminal, and the user can review its contents and submit approval or revision suggestions. This finalizes a feasible schedule that reflects the user's feedback.
[0555] Step 4:
[0556] The server generates a digital persona to represent the user in a virtual space. Inputs consist of approved schedule information and user instructions. The generated digital persona receives voice and text instructions from the user in real time and operates within the virtual environment. This allows the user to continue activities within the virtual environment without direct participation.
[0557] Step 5:
[0558] The server converts speech from digital characters and user voice input into text data in real time. The input is voice data from interactions in the virtual space. The output is the generated text data. This process allows the voice information to be saved in text format, making it easy to refer to later.
[0559] Step 6:
[0560] The server analyzes the text data and summarizes the key points. Text data is used as input, and summarized text is generated as output. This process allows users to concisely understand vast amounts of information and make important decisions quickly.
[0561] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0562] This invention provides a system that supports users in efficiently participating in meetings while also considering their emotions during communication. Specific embodiments for carrying out this invention are described below.
[0563] The server retrieves information from a database to manage user schedules and uses an artificial intelligence model to optimize meeting times. This process takes into account the user's preferences. The proposed schedule is then notified to the terminal by the server and can be reviewed or modified by the user. Once the user approves the schedule, the server stores it in the database and synchronizes it with the relevant applications.
[0564] For virtual meeting participation, the server generates an avatar, and the terminal reflects the user's voice and text input in real time. Furthermore, by incorporating an emotion engine, the server detects emotions from the user's facial expressions and voice patterns and reflects them in the avatar's expressions and movements. This allows the avatar to behave more naturally and human-like.
[0565] Audio data from meetings is converted into text data in real time by a server. The server analyzes the meeting content based on this text data and creates a summary. Furthermore, it analyzes the mood of the meeting using emotional information detected by an emotion engine and reports the impression and atmosphere to the user in a visualized form.
[0566] As a concrete example, consider a scenario where an employee participates in an online meeting. The server adjusts the user's schedule, ensuring a lunch break before setting up the afternoon meeting. During the meeting, an emotion engine determines whether the employee is relaxed or stressed, and appropriately changes the avatar's facial expressions accordingly. After the meeting, along with a summary of the meeting content, the server points out the periods when positive emotions were most prevalent, suggesting areas for improvement and successes for the user.
[0567] By implementing this invention, users can expect to maintain efficient and emotionally engaging communication, leading to improved work performance and results.
[0568] The following describes the processing flow.
[0569] Step 1:
[0570] The server retrieves schedule information from the database using the user ID. This information includes the user's schedule and desired break times.
[0571] Step 2:
[0572] The server uses an artificial intelligence model to calculate the optimal meeting time based on the acquired schedule information. In doing so, it takes into account the user's preferences.
[0573] Step 3:
[0574] The server sends the calculated schedule proposal to the terminal and notifies the user. The terminal displays this information and provides the user with options.
[0575] Step 4:
[0576] Users use their devices to review the proposed schedule and approve or modify it. User input is fed back to the server.
[0577] Step 5:
[0578] The server finalizes the schedule based on user feedback and saves it to the database. It also synchronizes with the calendar app as needed.
[0579] Step 6:
[0580] When a meeting begins, the server generates and places an avatar in the virtual space. The avatar acts as a representative of the user and prepares to receive the user's messages.
[0581] Step 7:
[0582] The device collects the user's voice and text input in real time and sends it to a server via the network. The server reflects the received content in the avatar and takes appropriate actions during the meeting.
[0583] Step 8:
[0584] The server uses an emotion engine to analyze the user's voice patterns and facial expressions, acquiring emotional information in real time. This information is reflected in the avatar's movements, naturally expressing the user's emotions.
[0585] Step 9:
[0586] During the meeting, the server converts audio information into text data in real time and saves it to the database. The converted data can be retrieved even after the meeting has ended.
[0587] Step 10:
[0588] After the meeting ends, the server analyzes the text data and sentiment information to generate a summary of the meeting content and sentiment. The generated summary and analysis report are provided to the user via their terminal.
[0589] Step 11:
[0590] Users can view the provided summary information and sentiment analysis reports to use for reviewing their work and preparing for future meetings.
[0591] (Example 2)
[0592] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0593] In today's busy environment, optimizing schedules and communicating with consideration for emotions are crucial for users to participate in meetings efficiently and smoothly. However, conventional schedule management systems and video conferencing tools have struggled to adequately reflect users' preferences and emotions when proposing schedules and communicating. Furthermore, it has been difficult to grasp the content of meetings and obtain meaningful feedback through emotional analysis during meetings. This invention aims to solve these problems and improve users' work efficiency and the quality of communication.
[0594] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0595] In this invention, the server includes means for acquiring schedule information by a device that collects and manages user schedule information, means for optimizing and proposing meeting times that reflect the user's desired conditions using a generated artificial intelligence model, and means for notifying the user's information device of the proposed meeting times and receiving approval or modification instructions. This enables the user to efficiently manage their schedule and participate in meetings optimally. Furthermore, by reflecting voice and text data onto a virtual entity representing the user and operating it in real time, it realizes natural communication that takes emotions into consideration. In addition, by converting voice data during meetings into text data in real time and providing a summary of the meeting content and the results of sentiment analysis during the meeting, the user can review their impression of the meeting and areas for improvement, and receive feedback that leads to improved work performance.
[0596] A "user" refers to an individual or member of an organization that uses the system, and is the entity that participates in meetings and manages schedules.
[0597] "Schedule information" refers to information related to the user's schedule, including data on the time and conditions of meetings, events, etc.
[0598] "Information equipment" refers to electronic devices used by users, including smartphones and personal computers.
[0599] A "data management device" is a device that provides means for storing, managing, and synchronizing data, and includes cloud servers and databases.
[0600] "Existence" refers to a digital character or avatar that acts on behalf of the user within a virtual realm.
[0601] The "generated artificial intelligence model" is a system composed of algorithms used for data analysis and prediction, and is particularly responsible for optimizing meeting times by taking user conditions into consideration.
[0602] "Real-time" refers to the characteristic that data processing and communication occur almost instantly, resulting in a state where the system responds to input without delay.
[0603] "Sentiment analysis" refers to the process of inferring and analyzing a user's emotional state based on their tone of voice and the content of their writing.
[0604] "Voice data conversion" refers to the process of converting voice information into text information, and is performed using speech recognition technology.
[0605] A "mood report" refers to information that visualizes the emotional state of participants and the atmosphere of a meeting, and is provided to users as feedback.
[0606] This invention provides a system that enables users to efficiently participate in meetings using information devices and to achieve emotionally sensitive communication. Specifically, a server collects user schedule information and manages it with a data management device. This allows the server to have a detailed understanding of the user's schedule information and use a generated artificial intelligence model to optimize meeting times based on the user's preferences.
[0607] The server notifies the user's information device of the optimized schedule and receives approval or modification instructions from the user. The modified schedule is stored by the server on a data management device and synchronized with the relevant application software. This allows the user to participate in meetings in a planned manner.
[0608] Furthermore, to support participation in virtual meetings, the server generates an avatar that acts as a proxy for the user. This avatar reflects the voice and text data transmitted from the user via the terminal and operates in real time. In addition, by incorporating an emotion analysis engine, the server detects the user's emotional state from their voice tone and speech content and immediately reflects it in the avatar's facial expressions. This function enables smooth communication that takes emotions into consideration.
[0609] During the meeting, the server converts audio data into text data in real time. This process uses speech recognition software. Based on the text data, the server summarizes the meeting content and analyzes the emotional information acquired during the meeting. The analysis results are provided to the user as a mood report, allowing the user to review the meeting's impression and key points and receive feedback.
[0610] For example, if a user has multiple meetings scheduled in the morning, the server uses a generative AI model to adjust morning and afternoon meeting times and suggest ways to ensure a lunch break. An example of a prompt message given to the generative AI model would be, "Optimize the schedule by adjusting afternoon meetings based on the user's preferences." This allows users to manage their schedules with less stress and achieve better work efficiency.
[0611] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0612] Step 1:
[0613] The server retrieves user schedule information from the data management device. The inputs used are the user's ID and associated past schedule information. Based on the retrieved schedule information, a data analysis algorithm extracts user trends and preferences, which are then provided to the AI model. The output generates the initial dataset necessary for schedule optimization.
[0614] Step 2:
[0615] The server uses the generated artificial intelligence model to optimize the schedule. In this process, the user's conditions and optimization goals are input to the generated AI model as prompts. The AI analyzes the input data and calculates the optimal time for meetings. As output, candidate meeting times are listed. As a concrete example, the prompt "Please schedule the meeting in the afternoon, avoiding lunchtime" is used.
[0616] Step 3:
[0617] The server notifies the user's information device of an optimized meeting schedule. The input is a list of meeting times suggested by a generative AI model. The server formats this as notification data and sends it to the user's terminal. As output, the suggested meeting schedule is displayed on the user's screen. The user reviews the presented dates and times and makes corrections as needed.
[0618] Step 4:
[0619] The user reviews the schedule and enters approval or modification information. The terminal receives the user's input and sends the modified schedule data back to the server. The input is the user's approval or modification instruction, and the output is the updated schedule data. This allows the user to manage schedules that meet their own criteria.
[0620] Step 5:
[0621] In preparation for a meeting in a virtual space, the server generates a user avatar. The inputs for this process are the user profile and past meeting data. The server configures the avatar's basic behavior and appearance, and outputs the data as a 3D model. During the meeting, the terminal reflects the user's voice and text onto the avatar, enabling real-time operation.
[0622] Step 6:
[0623] The server converts audio data into text data in real time. The input is an audio stream from a meeting, and the server uses speech recognition software to convert it into text. The output is the converted text data, which is used later for summarization and sentiment analysis. Specifically, it transcribes user speech into text, helping to visualize the meeting content.
[0624] Step 7:
[0625] The server summarizes meeting content and performs sentiment analysis based on text data. Input consists of real-time text data and elements for sentiment analysis. A data analysis algorithm processes this data to generate summarized meeting content and sentiment data. Output includes a visualized mood report and meeting summaries. Users can use this to assess the effectiveness of the meeting and improve future meetings.
[0626] (Application Example 2)
[0627] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0628] In today's commercial environment, online and offline portals are becoming increasingly integrated, but the user experience in virtual spaces is often limited. This results in a lack of deep interaction that considers user emotions, making it challenging to provide a service comparable to the experience in a physical store. Furthermore, there is a demand for flexible scheduling and personalized service that takes user emotions into account during meetings and business negotiations, but this is difficult to achieve with current technology.
[0629] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0630] In this invention, the server includes information processing means for collecting user schedule information, data analysis means for proposing a time that suits the user's preferences using an algorithm for optimizing meeting times based on the schedule information, and emotion analysis means for detecting the user's emotions and reflecting them in a corresponding digital representation. This enables a personalized virtual space experience that takes the user's emotions into consideration.
[0631] A "communication system" is a general term for technologies and devices used to send and receive data between information processing devices and terminals located in different places.
[0632] "Information processing means" refers to hardware or software functions for collecting, processing, and converting user data and information into a useful form.
[0633] A "data analysis tool" is a program or processing method that analyzes collected data and derives useful information according to a specific purpose.
[0634] "Emotional analysis means" refers to a technology or process that determines a user's psychological state and analyzes their emotions based on their voice, facial expressions, etc.
[0635] A "digital character" is a virtual entity that acts on behalf of the user on a computer and interacts with them.
[0636] "Generation means" refers to the process or device used to create new digital content based on specified data.
[0637] "Data conversion means" refers to a process or device that converts data in one format into another format, making it usable.
[0638] "Information display means" refers to technology or devices for visually providing data, results, or information to a user through an interface.
[0639] This invention is a system that enables effective communication in a virtual space while taking into account the user's emotions. The server collects the user's schedule information using information processing means and proposes the optimal meeting time using an algorithm based on the schedule information. To reflect the user's emotions, emotion analysis means analyzes the user's emotions from their voice and facial expressions and generates a corresponding digital representation. This digital representation is reflected in a digital character that acts as a proxy for the user in the virtual space, supporting interaction with the user.
[0640] During this process, the server uses Azure Kinect or equivalent emotion detection devices to acquire user emotion data in real time. Based on this data, a digital character is generated that uses speech synthesis services such as Amazon Polly to provide voice guidance and engage in natural communication with the user.
[0641] For example, if a user wants advice on a specific product in a virtual store, the system can determine from the user's cautious expression that they are hesitant about purchasing it. A digital character then gently encourages and supports the user by saying, "This product has been very well-received by many users. Why don't you give it a try?"
[0642] An example of a prompt is, "How can an AI model create a shopping guide that matches the user's emotions?" This prompt provides guidance to the generative AI model for generating appropriate guides and content.
[0643] In this way, the invention enables communication that takes emotions into account, leading to a more fulfilling experience for users.
[0644] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0645] Step 1:
[0646] The server collects schedule information from the user's smart glasses. As input, it retrieves data from the user's calendar or schedule. The server analyzes this data and prepares to suggest optimal meeting times based on the user's schedule patterns and preferences. The analyzed schedule information is then generated as output.
[0647] Step 2:
[0648] The server uses emotion analysis to acquire facial and voice data in real time from sensors installed in the user's smart glasses. The user's facial and voice data are provided as input. The server analyzes this data using an emotion detection algorithm to determine the user's emotional state. The user's emotional state is then generated as output.
[0649] Step 3:
[0650] The server determines the actions of the digital character in the virtual space based on the emotional state. The emotional state obtained in the previous step is used as input. This allows the digital character to reflect expressions and actions that correspond to the user's emotions, supporting interaction with the user. Action instructions for the digital character are generated as output.
[0651] Step 4:
[0652] The device controls a digital character and displays actions corresponding to the user's smart glasses. Input is provided as instructions for the digital character's actions from a server. The device uses a speech synthesis service such as Amazon Polly, allowing the digital character to make suggestions and explanations to the user in a natural voice. Output is generated as audio and visual feedback to the user.
[0653] Step 5:
[0654] The user reviews suggestions from the server and interacts with digital characters, then decides on actions such as making a purchase or attending a meeting. Input comes from information provided by the digital characters. Based on the user's actions and choices, the terminal records this information and sends it back to the server for further data analysis. Output is generated, which includes the user's selection information.
[0655] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0656] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0657] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0658] [Fourth Embodiment]
[0659] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0660] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0661] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0662] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0663] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0664] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0665] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0666] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0667] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0668] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0669] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0670] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0671] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0672] This invention provides a system that enables users to participate in meetings efficiently and reduce their workload. Specific embodiments for carrying out this invention are described below.
[0673] The server manages user schedule information by retrieving user-registered schedule data from a database. Based on the retrieved schedule information, it uses an artificial intelligence model to calculate the optimal time allocation, taking into account the user's desired meeting and break times. The server sends this calculation result to the terminal, and the schedule is finalized after the user confirms and approves it.
[0674] The user's device displays schedule suggestions received from the server as notifications and sends approvals or feedback to the server based on the user's input. Approved schedules are reflected in the user's calendar application.
[0675] For virtual conference participation, the server provides an avatar generation function, and the user's chosen avatar participates in the conference. The terminal detects the user's voice and text input in real time and reflects it in the avatar. In the virtual conference, the avatar speaks on behalf of the user and visualizes the conference content as needed.
[0676] Furthermore, the audio from the meeting is instantly converted into text data by the server and stored in a database. Based on this text data, the server analyzes the meeting content, extracts summaries and key topics, and provides them to the user via their terminal. This allows the user to quickly grasp the main points of the meeting and use them to their advantage in subsequent work.
[0677] As a concrete example, imagine a manager user who wants to adjust their schedule to ensure they have a lunch break. The AI then adjusts the schedule, allowing them to hold an afternoon meeting. In this scenario, the user participates in the meeting using an avatar, and after the meeting, they can quickly make decisions about subsequent processes based on a summary provided by the server.
[0678] By implementing this invention, users can reduce the burden of attending meetings, manage their time efficiently, and ultimately expect an improvement in overall work productivity.
[0679] The following describes the processing flow.
[0680] Step 1:
[0681] The server retrieves schedule information from the database using the user ID. This data includes the user's schedule and meeting information.
[0682] Step 2:
[0683] The server inputs the acquired schedule information into an artificial intelligence model and generates suggestions to optimize meeting and break times. The user's desired conditions (e.g., ensuring a lunch break) are taken into consideration.
[0684] Step 3:
[0685] The server sends the generated schedule proposal to the terminal and notifies the user. The terminal receives this information and displays it as a pop-up on the screen.
[0686] Step 4:
[0687] The user operates the terminal to review the proposed schedule and enter their approval or any necessary modifications. This input is sent to the server and used to finalize or further adjust the schedule.
[0688] Step 5:
[0689] The server records approved schedules in a database and synchronizes them with calendar applications, etc.
[0690] Step 6:
[0691] Once the meeting begins, the server generates and places the user's selected avatar within the virtual space. The avatar then prepares to reflect the user's statements.
[0692] Step 7:
[0693] The terminal receives user voice or text input in real time and sends it to the server via the network. The server then reflects this in the avatar's movements to facilitate communication within the meeting.
[0694] Step 8:
[0695] During the meeting, the server converts audio data into text, providing real-time transcription. This text is stored in a database and can be accessed even after the meeting ends.
[0696] Step 9:
[0697] Once the meeting ends, the server analyzes the text data and generates a summary. It extracts important topics and statements and sends them to the terminal in a visually easy-to-understand format.
[0698] Step 10:
[0699] Users can view summary information provided on their devices to review meeting content. Based on this, they can then adjust their work and make decisions.
[0700] (Example 1)
[0701] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0702] Traditional scheduling management systems had problems with efficient time allocation, as users had to individually set and adjust meeting times, which was time-consuming. Furthermore, in virtual meeting environments, incomplete information transmission due to a lack of real-world communication was a challenge. Additionally, the lack of features to effectively summarize and notify meeting information made it difficult to grasp the key points of the meeting.
[0703] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0704] In this invention, the server includes means for acquiring user schedule information from a data storage device, means for calculating an optimal time allocation according to the user's wishes using a generation AI model based on the acquired schedule information, means for notifying the user terminal of the calculated time allocation and receiving approval or correction feedback from the user, means for reflecting the approved time allocation in a calendar application program, means for generating a selected display image in a virtual space and having it operate as a user's proxy, means for reflecting voice and text information transmitted from the user via the terminal in the display image in real time, means for converting voice information collected during a meeting into text information in real time using voice processing means, information processing means for analyzing the text information and summarizing the meeting content, and means for providing the summarized information to the user through notification means. As a result, the user can efficiently manage their schedule, communicate smoothly in a virtual environment, and quickly grasp the key points of a meeting.
[0705] A "user" refers to an individual or group that has registered to use an information system.
[0706] "Schedule information" refers to a series of date, time, and event-related data registered and managed by the user.
[0707] A "data storage device" refers to a storage medium or device that can store and retrieve information.
[0708] A "generative AI model" is an artificial intelligence algorithm used to generate the optimal solution by performing complex calculations and reasoning based on human instructions.
[0709] "Time allocation" is a concept that refers to the optimal distribution of time according to specific activities or requirements.
[0710] A "terminal" is an electronic device used by a user to input and retrieve information, and usually refers to a computer or smartphone.
[0711] A "calendar application program" is a software application used for managing dates, times, and events.
[0712] A "virtual space" is a digital environment created by computer technology that users can interact with.
[0713] "Displayed image" refers to a visible substitute that acts as a representative of the user within a virtual space, such as an avatar.
[0714] "Speech processing means" refers to technologies and devices that analyze speech data and perform necessary conversions and analyses.
[0715] "Textual information" refers to information in text format converted from audio or other data.
[0716] "Information processing means" refers to technologies and devices that have the ability to analyze data and process it into a useful form according to a specific purpose.
[0717] "Notification means" refers to technologies and devices for informing a user of specific information visually or audibly.
[0718] This invention is a system that allows users to efficiently manage their schedules and effectively participate in meetings in a virtual environment. The following describes a specific implementation of this system.
[0719] The server first retrieves the user's schedule information from its data storage device. A common database management system can be used for this process. For example, the database stores the user's schedule information, and the server retrieves the necessary data using SQL queries.
[0720] Next, the server uses a generative AI model based on the acquired schedule information to calculate the optimal time allocation that meets the user's preferences. Various artificial intelligence platforms can be used as generative AI models in this process. In this process, the AI is instructed with a prompt such as, "Please tell me how to get the optimal meeting schedule suggested to the user," and the calculation results are obtained.
[0721] The user's device receives notifications of optimal time allocation sent from the server and displays them to the user. The device, using a smartphone or PC, provides an interface for the user to check, approve, or modify their schedule. This feedback is resent to the server, and the finally approved schedule is reflected in the calendar application program.
[0722] In virtual conference participation, the server generates a display image based on the user's selection. This display image acts as a proxy for the user in the virtual space and is implemented using common computer graphics technology. The user's terminal sends voice and text to the server via the terminal device, and the server reflects this in the display image in real time.
[0723] Furthermore, audio data from meetings is instantly converted into text information by a server using audio processing equipment. This text information is then analyzed by information processing equipment to extract summaries and key topics. The results of this analysis are provided to the user through notification systems, allowing them to quickly grasp the main points of the meeting. This entire system enables users to efficiently manage their schedules and effectively utilize the content of meetings.
[0724] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0725] Step 1:
[0726] The server retrieves the user's schedule information from the database. An SQL query is issued using the user's ID as input, and the corresponding schedule information is output from the database. Specifically, the server executes the query "SELECT FROM schedules WHERE user_id = User ID" and stores the retrieved data in its internal storage.
[0727] Step 2:
[0728] The server uses an AI model to calculate the optimal time allocation based on the acquired schedule information. The input is the schedule information and the prompt "Please suggest the optimal time allocation desired by the user." The output is the time allocation suggestion generated by the AI. The AI model performs internal calculations and generates a timetable based on the parameters to be terminated.
[0729] Step 3:
[0730] The server sends the generated time allocation to the user terminal. The input is the time allocation obtained in step 2, and the output is the notification to the user terminal. Specifically, the server uses a REST API to send an HTTP POST request to the user terminal, and the terminal receives the data and displays the notification on the screen.
[0731] Step 4:
[0732] The terminal displays proposals from the server to the user and accepts approval or modification from the user. The input is the time allocation proposal received from the server, and the output is the user's feedback. Specifically, the terminal screen displays "Approve" or "Modify" buttons, and after the user makes a selection, the feedback is sent to the server.
[0733] Step 5:
[0734] The server reflects the approved schedule in the calendar application program. The input is the user's approval feedback, and the output is the registration of the schedule in the calendar application program. Specifically, the server uses the calendar API and sends a request to add the schedule after authenticating with OAuth 2.0.
[0735] Step 6:
[0736] The server generates a display image of the user's selection for participation in a virtual meeting. The input is the user's selection information, and the output is the display image in the virtual space. The server uses a virtual environment engine to generate a display image based on the user's selection and reflect it in the virtual space.
[0737] Step 7:
[0738] The terminal sends voice and text from the user to the server, which is then reflected in the display in real time. Input is voice data and text information, and output is real-time reflection in a virtual space. The terminal collects information from the microphone and keyboard and sends it to the server sequentially.
[0739] Step 8:
[0740] The server converts audio information from a meeting into text, analyzes it, and generates a summary. The input is audio data, and the output is the summarized meeting content. Specifically, the server uses an audio processing engine to convert the audio into text, analyzes the text to extract key points, and generates a summary.
[0741] (Application Example 1)
[0742] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0743] In today's business environment, activities in virtual spaces and interactions using digital people are becoming increasingly important. However, existing technologies have not adequately addressed the challenges of user schedule management and the efficiency of actual conversations. This invention aims to provide a system that allows users to efficiently engage in activities in virtual spaces, reduce their workload, and conduct effective conversations.
[0744] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0745] In this invention, the server includes means for collecting user activity information managed by an information processing device, means for proposing a time allocation according to the user's wishes using an artificial intelligence model generated to optimize work time based on the activity information, and means for generating a digital person that acts as a proxy for the user in a virtual space. This enables the user to streamline activities in the virtual environment, reduce the user's workload, and improve the quality of interaction.
[0746] An "information processing device" is a computer device used to manage user activity information and schedules, and to control the entire system.
[0747] "Activity information" refers to data about various actions and schedules that users perform on a daily basis.
[0748] An "artificial intelligence model" is an AI algorithm used to analyze data and make optimal suggestions in order to optimize user activities.
[0749] "Time allocation" refers to the distribution of time to efficiently carry out activities and tasks according to the user's preferences.
[0750] A "virtual space" is a digital environment created by computers that is different from the real world.
[0751] A "digital character" is an avatar that acts on behalf of the user to perform various interactions in a virtual space.
[0752] "Real-time" refers to processing and responses occurring almost instantaneously, meaning that information is provided to the user without delay.
[0753] "Text data" refers to data obtained by converting spoken or oral information into text format.
[0754] A "summary" is a short, concise compilation of information or data, extracted from a large amount of information or data.
[0755] "Interaction" refers to the interaction or communication between a user and a system, or between users themselves.
[0756] This invention is a system that enables users to efficiently manage their activities and make the most of interactions in virtual space. Its main components include an information processing device, an artificial intelligence model, a digital person generation function, and a speech-to-text conversion function.
[0757] The server acts as an information processing device, collecting and managing user activity information. Considering user preferences and activity patterns, it optimizes work time using a generative AI model. To achieve this, it analyzes data using machine learning libraries such as TensorFlow and provides suggestions in a user-readable format. This allows users to create a more rational schedule.
[0758] The server also generates a digital persona to represent the user in the virtual space. This digital persona operates using real-time communication technologies such as WebRTC, reflecting the user's voice and text instructions. This allows users to communicate in a virtual environment without having to participate physically.
[0759] Furthermore, the server converts audio information into text data in real time. By using the Google Cloud Speech-to-Text API, the acquired audio information is efficiently converted into text and a summary of key topics is generated. This allows users to quickly grasp the main points of a conversation and use that information to aid in subsequent decision-making.
[0760] As a concrete example, suppose user A has a virtual meeting scheduled for the morning. The server proposes a pre-optimized schedule, and once user A confirms it, the digital persona joins the meeting and assists with the discussion. After the meeting ends, the server automatically summarizes the key points and provides them to user A. It is also possible to improve the quality of the digital persona's responses by using a prompt such as, "Suggest how the avatar should respond appropriately to the customer."
[0761] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0762] Step 1:
[0763] The server collects user activity information from a database. The input consists of schedules and daily activity data registered by the user. This collected information serves as the basis for optimization algorithms that are then processed. This prepares the system for more efficient management of user activity.
[0764] Step 2:
[0765] The server uses an artificial intelligence model to optimize work time based on the collected activity information. The collected activity information is used as input, and this information is analyzed to propose the optimal time allocation for each activity. This data processing is a necessary step to propose an efficient schedule to the user.
[0766] Step 3:
[0767] The server notifies the user's terminal of the optimized time allocation. The input is the optimized schedule generated by the artificial intelligence model. As output, the optimized schedule is displayed on the user's terminal, and the user can review its contents and submit approval or revision suggestions. This finalizes a feasible schedule that reflects the user's feedback.
[0768] Step 4:
[0769] The server generates a digital persona to represent the user in a virtual space. Inputs consist of approved schedule information and user instructions. The generated digital persona receives voice and text instructions from the user in real time and operates within the virtual environment. This allows the user to continue activities within the virtual environment without direct participation.
[0770] Step 5:
[0771] The server converts speech from digital characters and user voice input into text data in real time. The input is voice data from interactions in the virtual space. The output is the generated text data. This process allows the voice information to be saved in text format, making it easy to refer to later.
[0772] Step 6:
[0773] The server analyzes the text data and summarizes the key points. Text data is used as input, and summarized text is generated as output. This process allows users to concisely understand vast amounts of information and make important decisions quickly.
[0774] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0775] This invention provides a system that supports users in efficiently participating in meetings while also considering their emotions during communication. Specific embodiments for carrying out this invention are described below.
[0776] The server retrieves information from a database to manage user schedules and uses an artificial intelligence model to optimize meeting times. This process takes into account the user's preferences. The proposed schedule is then notified to the terminal by the server and can be reviewed or modified by the user. Once the user approves the schedule, the server stores it in the database and synchronizes it with the relevant applications.
[0777] For virtual meeting participation, the server generates an avatar, and the terminal reflects the user's voice and text input in real time. Furthermore, by incorporating an emotion engine, the server detects emotions from the user's facial expressions and voice patterns and reflects them in the avatar's expressions and movements. This allows the avatar to behave more naturally and human-like.
[0778] Audio data from meetings is converted into text data in real time by a server. The server analyzes the meeting content based on this text data and creates a summary. Furthermore, it analyzes the mood of the meeting using emotional information detected by an emotion engine and reports the impression and atmosphere to the user in a visualized form.
[0779] As a concrete example, consider a scenario where an employee participates in an online meeting. The server adjusts the user's schedule, ensuring a lunch break before setting up the afternoon meeting. During the meeting, an emotion engine determines whether the employee is relaxed or stressed, and appropriately changes the avatar's facial expressions accordingly. After the meeting, along with a summary of the meeting content, the server points out the periods when positive emotions were most prevalent, suggesting areas for improvement and successes for the user.
[0780] By implementing this invention, users can expect to maintain efficient and emotionally engaging communication, leading to improved work performance and results.
[0781] The following describes the processing flow.
[0782] Step 1:
[0783] The server retrieves schedule information from the database using the user ID. This information includes the user's schedule and desired break times.
[0784] Step 2:
[0785] The server uses an artificial intelligence model to calculate the optimal meeting time based on the acquired schedule information. In doing so, it takes into account the user's preferences.
[0786] Step 3:
[0787] The server sends the calculated schedule proposal to the terminal and notifies the user. The terminal displays this information and provides the user with options.
[0788] Step 4:
[0789] Users use their devices to review the proposed schedule and approve or modify it. User input is fed back to the server.
[0790] Step 5:
[0791] The server finalizes the schedule based on user feedback and saves it to the database. It also synchronizes with the calendar app as needed.
[0792] Step 6:
[0793] When a meeting begins, the server generates and places an avatar in the virtual space. The avatar acts as a representative of the user and prepares to receive the user's messages.
[0794] Step 7:
[0795] The device collects the user's voice and text input in real time and sends it to a server via the network. The server reflects the received content in the avatar and takes appropriate actions during the meeting.
[0796] Step 8:
[0797] The server uses an emotion engine to analyze the user's voice patterns and facial expressions, acquiring emotional information in real time. This information is reflected in the avatar's movements, naturally expressing the user's emotions.
[0798] Step 9:
[0799] During the meeting, the server converts audio information into text data in real time and saves it to the database. The converted data can be retrieved even after the meeting has ended.
[0800] Step 10:
[0801] After the meeting ends, the server analyzes the text data and sentiment information to generate a summary of the meeting content and sentiment. The generated summary and analysis report are provided to the user via their terminal.
[0802] Step 11:
[0803] Users can view the provided summary information and sentiment analysis reports to use for reviewing their work and preparing for future meetings.
[0804] (Example 2)
[0805] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0806] In today's busy environment, optimizing schedules and communicating with consideration for emotions are crucial for users to participate in meetings efficiently and smoothly. However, conventional schedule management systems and video conferencing tools have struggled to adequately reflect users' preferences and emotions when proposing schedules and communicating. Furthermore, it has been difficult to grasp the content of meetings and obtain meaningful feedback through emotional analysis during meetings. This invention aims to solve these problems and improve users' work efficiency and the quality of communication.
[0807] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0808] In this invention, the server includes means for acquiring schedule information by a device that collects and manages user schedule information, means for optimizing and proposing meeting times that reflect the user's desired conditions using a generated artificial intelligence model, and means for notifying the user's information device of the proposed meeting times and receiving approval or modification instructions. This enables the user to efficiently manage their schedule and participate in meetings optimally. Furthermore, by reflecting voice and text data onto a virtual entity representing the user and operating it in real time, it realizes natural communication that takes emotions into consideration. In addition, by converting voice data during meetings into text data in real time and providing a summary of the meeting content and the results of sentiment analysis during the meeting, the user can review their impression of the meeting and areas for improvement, and receive feedback that leads to improved work performance.
[0809] A "user" refers to an individual or member of an organization that uses the system, and is the entity that participates in meetings and manages schedules.
[0810] "Schedule information" refers to information related to the user's schedule, including data on the time and conditions of meetings, events, etc.
[0811] "Information equipment" refers to electronic devices used by users, including smartphones and personal computers.
[0812] A "data management device" is a device that provides means for storing, managing, and synchronizing data, and includes cloud servers and databases.
[0813] "Existence" refers to a digital character or avatar that acts on behalf of the user within a virtual realm.
[0814] The "generated artificial intelligence model" is a system composed of algorithms used for data analysis and prediction, and is particularly responsible for optimizing meeting times by taking user conditions into consideration.
[0815] "Real-time" refers to the characteristic that data processing and communication occur almost instantly, resulting in a state where the system responds to input without delay.
[0816] "Sentiment analysis" refers to the process of inferring and analyzing a user's emotional state based on their tone of voice and the content of their writing.
[0817] "Voice data conversion" refers to the process of converting voice information into text information, and is performed using speech recognition technology.
[0818] A "mood report" refers to information that visualizes the emotional state of participants and the atmosphere of a meeting, and is provided to users as feedback.
[0819] This invention provides a system that enables users to efficiently participate in meetings using information devices and to achieve emotionally sensitive communication. Specifically, a server collects user schedule information and manages it with a data management device. This allows the server to have a detailed understanding of the user's schedule information and use a generated artificial intelligence model to optimize meeting times based on the user's preferences.
[0820] The server notifies the user's information device of the optimized schedule and receives approval or modification instructions from the user. The modified schedule is stored by the server on a data management device and synchronized with the relevant application software. This allows the user to participate in meetings in a planned manner.
[0821] Furthermore, to support participation in virtual meetings, the server generates an avatar that acts as a proxy for the user. This avatar reflects the voice and text data transmitted from the user via the terminal and operates in real time. In addition, by incorporating an emotion analysis engine, the server detects the user's emotional state from their voice tone and speech content and immediately reflects it in the avatar's facial expressions. This function enables smooth communication that takes emotions into consideration.
[0822] During the meeting, the server converts audio data into text data in real time. This process uses speech recognition software. Based on the text data, the server summarizes the meeting content and analyzes the emotional information acquired during the meeting. The analysis results are provided to the user as a mood report, allowing the user to review the meeting's impression and key points and receive feedback.
[0823] For example, if a user has multiple meetings scheduled in the morning, the server uses a generative AI model to adjust morning and afternoon meeting times and suggest ways to ensure a lunch break. An example of a prompt message given to the generative AI model would be, "Optimize the schedule by adjusting afternoon meetings based on the user's preferences." This allows users to manage their schedules with less stress and achieve better work efficiency.
[0824] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0825] Step 1:
[0826] The server retrieves user schedule information from the data management device. The inputs used are the user's ID and associated past schedule information. Based on the retrieved schedule information, a data analysis algorithm extracts user trends and preferences, which are then provided to the AI model. The output generates the initial dataset necessary for schedule optimization.
[0827] Step 2:
[0828] The server uses the generated artificial intelligence model to optimize the schedule. In this process, the user's conditions and optimization goals are input to the generated AI model as prompts. The AI analyzes the input data and calculates the optimal time for meetings. As output, candidate meeting times are listed. As a concrete example, the prompt "Please schedule the meeting in the afternoon, avoiding lunchtime" is used.
[0829] Step 3:
[0830] The server notifies the user's information device of an optimized meeting schedule. The input is a list of meeting times suggested by a generative AI model. The server formats this as notification data and sends it to the user's terminal. As output, the suggested meeting schedule is displayed on the user's screen. The user reviews the presented dates and times and makes corrections as needed.
[0831] Step 4:
[0832] The user reviews the schedule and enters approval or modification information. The terminal receives the user's input and sends the modified schedule data back to the server. The input is the user's approval or modification instruction, and the output is the updated schedule data. This allows the user to manage schedules that meet their own criteria.
[0833] Step 5:
[0834] In preparation for a meeting in a virtual space, the server generates a user avatar. The inputs for this process are the user profile and past meeting data. The server configures the avatar's basic behavior and appearance, and outputs the data as a 3D model. During the meeting, the terminal reflects the user's voice and text onto the avatar, enabling real-time operation.
[0835] Step 6:
[0836] The server converts audio data into text data in real time. The input is an audio stream from a meeting, and the server uses speech recognition software to convert it into text. The output is the converted text data, which is used later for summarization and sentiment analysis. Specifically, it transcribes user speech into text, helping to visualize the meeting content.
[0837] Step 7:
[0838] The server summarizes meeting content and performs sentiment analysis based on text data. Input consists of real-time text data and elements for sentiment analysis. A data analysis algorithm processes this data to generate summarized meeting content and sentiment data. Output includes a visualized mood report and meeting summaries. Users can use this to assess the effectiveness of the meeting and improve future meetings.
[0839] (Application Example 2)
[0840] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0841] In today's commercial environment, online and offline portals are becoming increasingly integrated, but the user experience in virtual spaces is often limited. This results in a lack of deep interaction that considers user emotions, making it challenging to provide a service comparable to the experience in a physical store. Furthermore, there is a demand for flexible scheduling and personalized service that takes user emotions into account during meetings and business negotiations, but this is difficult to achieve with current technology.
[0842] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0843] In this invention, the server includes information processing means for collecting user schedule information, data analysis means for proposing a time that suits the user's preferences using an algorithm for optimizing meeting times based on the schedule information, and emotion analysis means for detecting the user's emotions and reflecting them in a corresponding digital representation. This enables a personalized virtual space experience that takes the user's emotions into consideration.
[0844] A "communication system" is a general term for technologies and devices used to send and receive data between information processing devices and terminals located in different places.
[0845] "Information processing means" refers to hardware or software functions for collecting, processing, and converting user data and information into a useful form.
[0846] A "data analysis tool" is a program or processing method that analyzes collected data and derives useful information according to a specific purpose.
[0847] "Emotional analysis means" refers to a technology or process that determines a user's psychological state and analyzes their emotions based on their voice, facial expressions, etc.
[0848] A "digital character" is a virtual entity that acts on behalf of the user on a computer and interacts with them.
[0849] "Generation means" refers to the process or device used to create new digital content based on specified data.
[0850] "Data conversion means" refers to a process or device that converts data in one format into another format, making it usable.
[0851] "Information display means" refers to technology or devices for visually providing data, results, or information to a user through an interface.
[0852] This invention is a system that enables effective communication in a virtual space while taking into account the user's emotions. The server collects the user's schedule information using information processing means and proposes the optimal meeting time using an algorithm based on the schedule information. To reflect the user's emotions, emotion analysis means analyzes the user's emotions from their voice and facial expressions and generates a corresponding digital representation. This digital representation is reflected in a digital character that acts as a proxy for the user in the virtual space, supporting interaction with the user.
[0853] During this process, the server uses Azure Kinect or equivalent emotion detection devices to acquire user emotion data in real time. Based on this data, a digital character is generated that uses speech synthesis services such as Amazon Polly to provide voice guidance and engage in natural communication with the user.
[0854] For example, if a user wants advice on a specific product in a virtual store, the system can determine from the user's cautious expression that they are hesitant about purchasing it. A digital character then gently encourages and supports the user by saying, "This product has been very well-received by many users. Why don't you give it a try?"
[0855] An example of a prompt is, "How can an AI model create a shopping guide that matches the user's emotions?" This prompt provides guidance to the generative AI model for generating appropriate guides and content.
[0856] In this way, the invention enables communication that takes emotions into account, leading to a more fulfilling experience for users.
[0857] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0858] Step 1:
[0859] The server collects schedule information from the user's smart glasses. As input, it retrieves data from the user's calendar or schedule. The server analyzes this data and prepares to suggest optimal meeting times based on the user's schedule patterns and preferences. The analyzed schedule information is then generated as output.
[0860] Step 2:
[0861] The server uses emotion analysis to acquire facial and voice data in real time from sensors installed in the user's smart glasses. The user's facial and voice data are provided as input. The server analyzes this data using an emotion detection algorithm to determine the user's emotional state. The user's emotional state is then generated as output.
[0862] Step 3:
[0863] The server determines the actions of the digital character in the virtual space based on the emotional state. The emotional state obtained in the previous step is used as input. This allows the digital character to reflect expressions and actions that correspond to the user's emotions, supporting interaction with the user. Action instructions for the digital character are generated as output.
[0864] Step 4:
[0865] The device controls a digital character and displays actions corresponding to the user's smart glasses. Input is provided as instructions for the digital character's actions from a server. The device uses a speech synthesis service such as Amazon Polly, allowing the digital character to make suggestions and explanations to the user in a natural voice. Output is generated as audio and visual feedback to the user.
[0866] Step 5:
[0867] The user reviews suggestions from the server and interacts with digital characters, then decides on actions such as making a purchase or attending a meeting. Input comes from information provided by the digital characters. Based on the user's actions and choices, the terminal records this information and sends it back to the server for further data analysis. Output is generated, which includes the user's selection information.
[0868] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0869] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0870] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0871] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0872] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0873] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0874] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0875] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0876] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0877] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0878] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0879] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0880] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0881] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0882] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0883] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0884] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0885] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0886] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0887] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0888] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0889] The following is further disclosed regarding the embodiments described above.
[0890] (Claim 1)
[0891] A means for collecting user schedule information managed by an information processing device,
[0892] A means for proposing meeting times that meet the user's preferences, using an artificial intelligence model generated to optimize meeting times based on the aforementioned schedule information,
[0893] A means of notifying the user's terminal of the proposed meeting time and receiving feedback from the user for approval or modification,
[0894] A schedule management system that includes this.
[0895] (Claim 2)
[0896] A means of generating an avatar that acts as a proxy for the user in a virtual space,
[0897] A means for making the aforementioned avatar operate in real time by reflecting voice and text information sent from the user,
[0898] The system according to claim 1, including the following:
[0899] (Claim 3)
[0900] A speech-to-text conversion method that converts audio information acquired during a meeting into text data in real time,
[0901] A means for analyzing the aforementioned text data and summarizing the meeting content,
[0902] Means for providing the user with the summarized meeting content,
[0903] The system according to claim 1, including the following:
[0904] "Example 1"
[0905] (Claim 1)
[0906] A means of obtaining user schedule information from a data storage device,
[0907] A method for calculating the optimal time allocation according to the user's wishes using a generated AI model based on acquired schedule information,
[0908] A means of notifying the user terminal of the calculated time allocation and receiving approval or correction feedback from the user,
[0909] A means to reflect the approved time allocation in the calendar application program,
[0910] A schedule management system that includes this feature.
[0911] (Claim 2)
[0912] A means for generating a selected display image in a virtual space and making it act as a user's proxy,
[0913] A means for reflecting voice and text information transmitted from the user via a terminal onto the display in real time,
[0914] The system according to claim 1, including the following:
[0915] (Claim 3)
[0916] A means for converting audio information collected during a meeting into text information in real time using audio processing equipment,
[0917] Information processing means for analyzing textual information and summarizing meeting content,
[0918] A means of providing summarized information to the user through a notification means,
[0919] The system according to claim 1, including the following:
[0920] "Application Example 1"
[0921] (Claim 1)
[0922] A means for collecting user activity information managed by an information processing device,
[0923] A means for proposing a time allocation according to the user's wishes, using an artificial intelligence model generated to optimize work time based on the aforementioned activity information,
[0924] A means of notifying the user's device of the proposed time and receiving feedback from the user for approval or modification,
[0925] A means for generating a digital person that acts as a proxy for a user in a virtual space,
[0926] A means for making the aforementioned digital person operate in real time by reflecting voice and text information transmitted from the user,
[0927] A means of converting audio information acquired during a conversation into text data in real time,
[0928] A means for analyzing the aforementioned text data and summarizing the content of the dialogue,
[0929] Means for providing the summarized dialogue content to the user,
[0930] A system that includes this.
[0931] (Claim 2)
[0932] The system according to claim 1, comprising means for recording information exchanged between a user's digital persona in a virtual space and generating log data for later analysis.
[0933] (Claim 3)
[0934] The system according to claim 1, comprising means for adjusting user activity in real time based on a newly generated AI model to optimize user activity.
[0935] "Example 2 of combining an emotion engine"
[0936] (Claim 1)
[0937] A means for acquiring schedule information by a device that collects and manages user schedule information,
[0938] A means for optimizing and proposing meeting times that reflect the user's desired conditions using an artificial intelligence model generated based on the aforementioned schedule information,
[0939] A means for notifying the user's information device of the proposed meeting time and receiving approval or modification instructions from the user,
[0940] A means for saving approved schedules to a data management device and synchronizing them with related application software,
[0941] A system that includes this.
[0942] (Claim 2)
[0943] A means of creating an entity that acts as a proxy for the user within a virtual domain,
[0944] The aforementioned entity is operated in real time by reflecting voice and text data from the user, and includes means for detecting the user's emotions and reflecting them in the aforementioned entity.
[0945] The system according to claim 1.
[0946] (Claim 3)
[0947] A voice data conversion method that converts audio information collected during a meeting into text data in real time,
[0948] A means for analyzing and summarizing the meeting content based on the aforementioned text data,
[0949] The means of providing the user with a mood report including the summarized meeting content and the results of the sentiment analysis during the meeting, and includes means for providing educational feedback to the user.
[0950] The system according to claim 1.
[0951] "Application example 2 when combining with an emotional engine"
[0952] (Claim 1)
[0953] Information processing means for collecting user schedule information,
[0954] A data analysis means that proposes meeting times according to the user's preferences, using an algorithm generated to optimize meeting times based on the aforementioned schedule information,
[0955] Information processing means that notifies the user's terminal of the proposed meeting time and receives feedback from the user for approval or modification,
[0956] A means of analyzing user emotions and reflecting them in corresponding digital representations,
[0957] A communication system that includes this.
[0958] (Claim 2)
[0959] A means for generating a digital character that acts as a proxy for the user in a virtual space,
[0960] Information processing means that makes the aforementioned digital character operate in real time by reflecting voice and text information transmitted from the user,
[0961] Control means for the aforementioned digital character to behave in accordance with the user's emotions,
[0962] A communication system according to claim 1, including the following:
[0963] (Claim 3)
[0964] A data conversion method that converts audio information acquired during a meeting into text data in real time,
[0965] A data processing means for analyzing the aforementioned text data and summarizing the meeting content,
[0966] Information display means that provides the user with analytical information based on the summarized meeting content and the participants' emotions,
[0967] A communication system according to claim 1, including the following: [Explanation of Symbols]
[0968] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for collecting user schedule information managed by an information processing device, A means for proposing meeting times that meet the user's preferences, using an artificial intelligence model generated to optimize meeting times based on the aforementioned schedule information, A means of notifying the user's terminal of the proposed meeting time and receiving feedback from the user for approval or modification, A schedule management system that includes this.
2. A means of generating an avatar that acts as a proxy for the user in a virtual space, A means for making the aforementioned avatar operate in real time by reflecting voice and text information sent from the user, The system according to claim 1, including the following:
3. A speech-to-text conversion method that converts audio information acquired during a meeting into text data in real time, A means for analyzing the aforementioned text data and summarizing the meeting content, Means for providing the user with the summarized meeting content, The system according to claim 1, including the following:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A