system

The system addresses inefficiencies in meetings by generating intelligent agents to conduct virtual discussions and summarize results, improving meeting efficiency and reducing the need for physical meetings.

JP2026070112APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

Smart Images

  • Figure 2026070112000001_ABST
    Figure 2026070112000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of inputting information, and a means of receiving the meeting agenda and participants' opinions as data, A means for analyzing the aforementioned data and generating a surrogate intelligent agent appropriate to the participant's position, A means of conducting a virtual meeting among the generated intelligent agents and obtaining the results of the discussion, A means of summarizing the results of a virtual meeting and providing them as a summary, Based on the summary provided, a means to determine whether an actual meeting is necessary, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern enterprises and organizations, there is a problem that a lot of business time is spent on meetings. In particular, in meetings attended by a large number of people, it may take an excessive amount of time to reconcile opinions. In addition, there are many meetings held only for the purpose of obtaining comments, confirmations, and approvals between departments, which reduces the efficiency of the entire business. There is a need to reduce the time required for such meetings and improve the efficiency of meetings.

Means for Solving the Problems

[0005] This invention provides a system that takes information as input and receives meeting agendas and participants' opinions as data. This system has the function of analyzing the received data and generating a surrogate intelligent agent appropriate to the participant's position. The generated intelligent agent conducts a virtual meeting, obtains the results, and summarizes them. Furthermore, by providing this summarized summary, it is possible to determine whether an actual meeting is necessary. In this way, the aim is to reduce meeting time and improve work efficiency.

[0006] "Means of inputting information" refers to a system where users provide the system with meeting agendas and participants' opinions as data.

[0007] "Means of analyzing and generating surrogate intelligent agents tailored to the participants' positions" refers to the process of analyzing received data and creating AI agents that mimic the opinions and positions of each participant.

[0008] "A means of conducting virtual meetings between intelligent agents and obtaining the results of the discussions" refers to a process in which generated AI agents exchange opinions in a virtual space, and information such as points of agreement and differences is extracted as a result.

[0009] "A means of summarizing and providing a summary" refers to a function that condenses the information obtained in a virtual meeting, organizes the main points in an easy-to-understand manner, and presents them to the user.

[0010] "Methods for determining whether an actual meeting is necessary" refers to the process of evaluating whether an additional meeting is needed in the real world, based on the generated summary. [Brief explanation of the drawing]

[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0013] First, let's explain the terminology used in the following explanation.

[0014] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0015] <0OO0092>In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0016] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0017] In the following embodiments, the labeled communication I / F (Interface) is an interface that includes a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0019] [First Embodiment]

[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0032] In order to implement the present invention, it is necessary to construct a system equipped with fundamental components. The main components of this system include a server, a terminal, and user interfaces. Specific embodiments are described below.

[0033] First, the user inputs the meeting agenda and participants' opinions using a terminal. This data is sent to the server via the terminal. The server receives this data and generates intelligent agents tailored to each participant's position and opinions. The generated AI agents have the ability to mimic each participant's perspective and conduct a virtual meeting.

[0034] The server holds virtual meetings among intelligent agents and gathers information through discussion. The server summarizes the results obtained during this process, creating a human-readable summary. The server then sends this summary to the terminal, which the user can view.

[0035] Based on the provided summary, users can determine whether an actual meeting needs to be held, enabling efficient decision-making. Consider a project management meeting as a concrete example. The user collects opinions from participants in advance and sends them to the server. The server analyzes this data and generates intelligent agents for each participant. In the virtual meeting, the agents exchange opinions and summarize the results. Finally, a statistical analysis report or a summary in presentation format is provided to the user, reducing the need for subsequent discussions.

[0036] This embodiment allows for the smooth preparation and operation of meetings involving large numbers of people, enabling efficient use of work time. The use of this system is also expected to reduce the overall length of meetings.

[0037] The following describes the processing flow.

[0038] Step 1:

[0039] The user inputs the meeting agenda and participants' opinions into a terminal. The terminal converts this into the system's input data format and sends it to the server.

[0040] Step 2:

[0041] The server analyzes the data received from the terminal. Based on the results of the data analysis, it generates intelligent agents tailored to each participant's opinions and perspectives.

[0042] Step 3:

[0043] The server hosts a virtual meeting among the generated intelligent agents. Each agent presents their opinion on the agenda and the discussion progresses through interaction with other agents.

[0044] Step 4:

[0045] The server extracts and summarizes the conclusions and key discussion points obtained during the virtual meeting. The summarized summary is organized to include key points, agreements, and, if necessary, further considerations.

[0046] Step 5:

[0047] The server sends the generated summary to the terminal. The user reviews the summary via the terminal and decides whether an actual meeting needs to be held. If the summary is sufficient, the meeting can be omitted.

[0048] (Example 1)

[0049] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0050] Currently, preparing and running large-scale meetings requires a tremendous amount of time and effort. In particular, gathering participants' opinions in advance and conducting efficient discussions based on those opinions is difficult. Traditional methods often result in insufficient information sharing and consensus building before the meeting, leading to prolonged meetings. Therefore, there is a need for efficient methods of information gathering and discussion beforehand.

[0051] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0052] In this invention, the server includes a device for inputting information, a device means for receiving topics related to the meeting and participants' opinions as digital data, a device means for analyzing the digital data and generating artificial intelligence agents according to the roles of the participants, and a device means for conducting a virtual meeting among the generated artificial intelligence agents and obtaining the results of the discussion. This makes it possible to integrate participants' opinions in advance through the virtual meeting and improve the efficiency of the meeting.

[0053] "A device for inputting information" refers to a device or interface for users to input the topic of a meeting or the opinions of participants.

[0054] A "device for receiving meeting-related topics and participants' opinions as digital data" refers to a device that receives meeting content and participants' opinions in digital format and makes them analyzable.

[0055] A "device that analyzes digital data and generates artificial intelligence agents according to the roles of the participants" refers to a device that creates artificial intelligence agents based on the received digital data, according to the position and opinions of each participant.

[0056] A "device for conducting virtual meetings among generated artificial intelligence agents and obtaining the results of discussions" refers to a device that allows multiple artificial intelligence agents to engage in discussions in a virtual space and obtain the resulting information.

[0057] A "device that provides data as a summary" refers to a device that summarizes and provides the results of a virtual meeting in a format that is easy for humans to understand.

[0058] "A device that determines whether an actual meeting is necessary based on the provided summary data" refers to a device that uses summarized information to determine whether or not a physical meeting needs to be held.

[0059] In order to implement the present invention, it is necessary to construct a system in which each component works in coordination. Specifically, it is important that the server, terminal, and user interface work together.

[0060] Users will input meeting content and participants' opinions using their devices. Personal computers and tablet devices can be used as this interface. The software is expected to include common communication tools such as messaging platforms and document creation tools.

[0061] The terminal sends the entered data to the server. The data is transmitted via a protocol such as HTTPS, ensuring the security of the communication. The data format should preferably be one that can be efficiently processed by the server, such as JSON or XML.

[0062] The server can utilize data processing libraries to analyze the received data. Examples include Python's Pandas and NumPy. Based on the results of this analysis, an artificial intelligence agent is generated using a generative AI model. For the generative AI, a model excelling in natural language processing, such as a large-scale language model, is used.

[0063] The generated artificial intelligence agent mimics the opinions of the participants and conducts a virtual meeting. The server manages this virtual meeting and records the content of the discussion. The information obtained from the discussion is summarized using NLP techniques. The specific techniques used here include the OpenAI® API and other natural language processing frameworks.

[0064] Finally, the generated summary is sent to the device and displayed on the user's device. The user uses this summary as a reference to decide whether or not to hold an actual meeting, if necessary. Throughout this entire process, meeting preparation is made more efficient, and users can make important decisions in a short amount of time.

[0065] As a concrete example, consider a meeting regarding the market launch strategy for a new product. Users collect participants' opinions using their devices, and these opinions are sent to a server. The server builds an AI agent based on each participant's opinion and finds a harmonious solution through a virtual discussion.

[0066] An example of a prompt to input into the generating AI model would be: "Simulate a virtual meeting based on the following meeting agenda and participant opinions, and summarize the results. Agenda: New product launch strategy. Participant opinions: Participant A…, Participant B…"

[0067] Thus, the present invention enables efficient meeting management and allows for the virtual integration of diverse perspectives from participants.

[0068] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0069] Step 1:

[0070] Users input meeting agenda items and participants' opinions into a terminal. Personal computers or tablet devices are used as input interfaces. The input data includes detailed opinions from each participant, thereby collecting the perspectives of the meeting attendees.

[0071] Step 2:

[0072] The terminal converts the input data into JSON format and sends it to the server using the HTTPS protocol. This ensures data security and efficient delivery to the server. The output is data in a format that can be processed by the server.

[0073] Step 3:

[0074] The server parses the received JSON data. This parsing uses data processing libraries such as Python's Pandas and NumPy. Specifically, it organizes and structures each participant's opinion as text data. The output at this stage is a parsed data frame.

[0075] Step 4:

[0076] The server invokes a generative AI model based on the analyzed data to generate artificial intelligence agents that mimic the perspectives of each participant. A large-scale language model is used as the generative AI model. For this generation, prompt sentences reflecting the opinions of each participant are prepared and input into the AI ​​model. The output is multiple artificial intelligence agents.

[0077] Step 5:

[0078] The server initiates a virtual meeting among the generated artificial intelligence agents. Each agent exchanges opinions according to their respective roles and the discussion progresses. Information is collected and stored in real time, and a discussion log is generated as output.

[0079] Step 6:

[0080] The server summarizes the logs obtained from the virtual meeting. Natural language processing (NLP) techniques are used to extract the main points, agreements, and disagreements of the discussion. The output is a human-readable text summary.

[0081] Step 7:

[0082] The server sends the generated summary to the terminal. Again, it is transmitted securely using the HTTPS protocol. The user can receive this summary on their terminal and understand its content. The output is a document or presentation material containing the summary.

[0083] Step 8:

[0084] Users use the provided summary to decide whether or not to actually hold a meeting. This allows for prior consensus building, reduces actual meeting time, and makes decision-making more efficient. The output is a meeting schedule based on the user's decision.

[0085] (Application Example 1)

[0086] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0087] Modern factory environments demand increased work efficiency and smoother operations. However, this presents challenges due to the need for communication between multiple pieces of equipment and participants, which can be time-consuming to coordinate. To address this, it is necessary to effectively incorporate participants' opinions, generate optimal work schedules, and support rapid decision-making on-site.

[0088] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0089] In this invention, the server includes a device for inputting information, a device for receiving the meeting topic and participants' opinions as data, a device for analyzing the data and generating a surrogate intelligent agent appropriate to the participants' positions, and a device for inputting on-site environmental data and analyzing the operating status of the device. This enables efficient imitation of each participant's perspective and allows for rapid and accurate decision-making on-site.

[0090] A "device for inputting information" is a device that allows users to collect various types of data and input them into a system.

[0091] The "topic of the meeting" refers to the specific topic or theme that will be discussed.

[0092] "Participant opinions" refer to the individual suggestions and views that each participant holds during the meeting.

[0093] A "device that receives data" is a device that collects information and data from external sources and incorporates it into a system in an appropriate format.

[0094] An "analytical device" is a device that analyzes input data and extracts the underlying meanings and patterns.

[0095] A "proxy intelligence agent" is a program that aims to mimic the position and opinions of a specific participant and virtually fulfill that role.

[0096] "Virtual dialogue" refers to the exchange of opinions and communication that takes place between agents in a simulated manner.

[0097] A "device for obtaining the results of a discussion" is a device that receives the results of virtual dialogues between agents and aggregates their contents.

[0098] The "device provided as a report" is a device that summarizes the results of a virtual meeting and provides feedback to the user in an easy-to-understand format.

[0099] "On-site environmental data" refers to data that describes the conditions and circumstances at the work site.

[0100] "Equipment operating status" refers to information indicating how production equipment and machinery are currently working.

[0101] An "optimal work schedule" refers to a timetable or plan formulated to efficiently and effectively carry out tasks or projects.

[0102] This invention provides a system for achieving efficient management in a factory. The system aims to digitize on-site work conditions and support rapid decision-making.

[0103] The server is connected to an information input device and receives meeting topics and participants' opinions as data. Terminal devices (e.g., smart glasses) input environmental data collected from the field in real time. The data entered on the terminal is sent to the server. The server uses Python to analyze the data and generates a surrogate intelligent agent that mimics the perspective of the participants.

[0104] The generated intelligent agent uses embellished prompts to conduct a virtual dialogue based on the generated AI model. The server dynamically retrieves the results of this virtual dialogue and summarizes them as a report. The report is provided through a terminal device to facilitate appropriate responses on-site.

[0105] As a concrete example, if a delay occurs on a factory production line, this system analyzes the operating status of each piece of equipment based on data collected from sensors. It then generates an optimal work schedule and reports it to the manager. This enables swift countermeasures.

[0106] An example of a prompt message is, "Use sensors to detect the current status of the factory equipment and create a schedule to optimize the next maintenance." In this way, the system supports efficient business operations.

[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0108] Step 1:

[0109] The terminal receives on-site environmental data and comments from the user. The entered data includes on-site sensor information and comments from the administrator. This data is then sent from the terminal to the server.

[0110] Step 2:

[0111] The server analyzes the received data. Using Python, it processes environmental data, including the operating status and signs of anomalies of each device, and extracts patterns and trends. This analysis provides the foundational data for generating intelligent agents.

[0112] Step 3:

[0113] The server generates surrogate intelligent agents using a generative AI model based on preprocessed data. It creates prompt statements and instructs each agent to mimic individual opinions and perspectives. This step takes into account past discussion history and information related to specific topics.

[0114] Step 4:

[0115] The server facilitates virtual dialogue between the generated intelligent agents. Each agent exchanges opinions and advances the discussion based on pre-configured prompts. This process facilitates opinion coordination within the virtual space.

[0116] Step 5:

[0117] The server retrieves the results of the virtual dialogue and summarizes key information and insights. It performs data analysis and extracts key points of agreement and areas requiring further discussion as a summary.

[0118] Step 6:

[0119] The server sends the generated summary to the terminal. The user reviews this summary and makes the best decision based on the situation on site. This final result forms the basis for determining the next action.

[0120] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0121] This invention is a system for improving the efficiency of meetings, and has a configuration that combines a user emotion engine that recognizes user emotions. In this system, the server, terminals, and the emotion analysis engine work together in coordination.

[0122] In terms of usage, users input meeting agenda items and participants' opinions using a terminal. This information is sent to the server, which then generates intelligent agents for each participant based on the received information. During the generation of intelligent agents, the user's emotional data, analyzed by the emotion engine, is also taken into consideration. Specifically, emotional data is obtained from voice and text to understand the emotional responses of the participants.

[0123] The server virtually conducts meetings between intelligent agents, utilizing an emotion engine to gather emotional information and use it to guide the discussion and determine its importance. The intelligent agents can offer opinions and adjust the discussion based on emotional triggers. In this way, the user's emotional state is woven into the flow of the meeting, enabling a more refined and nuanced discussion.

[0124] The results of the virtual meeting are summarized by the server and sent to the terminal as a summary, including emotional feedback. By referring to this summary, users can evaluate the necessity of the actual meeting and make efficient decisions. For example, in a product development meeting, if concerns from the technical department are accompanied by strong emotional reactions, this point will be highlighted in the summary, and the user can hold a separate meeting specifically for that issue.

[0125] This configuration enhances the effectiveness of meetings and provides a new approach for users to advance discussions more quickly and accurately. The feedback system utilizing the emotion engine goes beyond mere opinion gathering and can also take into account the inner state of participants, thus enabling more comprehensive meeting management.

[0126] The following describes the processing flow.

[0127] Step 1:

[0128] The user uses a terminal to input the meeting agenda and the opinions of the participating members. This information may include data in voice or text format. The terminal prepares to send the entered information to the server.

[0129] Step 2:

[0130] The server analyzes the data received from the terminal. This data analysis includes not only the opinions provided by the user, but also the acquisition of various sentiment indicators. The server uses a sentiment engine to recognize the user's emotions from voice and text and add them to the data.

[0131] Step 3:

[0132] The server generates intelligent agents for each participant based on the analysis results. During this generation process, the acquired emotional information is taken into consideration, and each agent is configured to mimic the emotions and perspectives of the participants.

[0133] Step 4:

[0134] The server conducts virtual meetings among the intelligent agents. In these virtual meetings, each agent expresses their opinions and positions and interacts with other agents. Based on emotional information, agents adjust the timing of their counterarguments and endorsements to optimize the flow of the discussion.

[0135] Step 5:

[0136] The server summarizes the results of the virtual meeting. The summarization process extracts key points based on main points of agreement, disagreement, and emotional responses. Emotional feedback is also reflected in the summary, clearly identifying points that require adjustment.

[0137] Step 6:

[0138] The server sends the generated summary to the terminal. The user reviews the summary on the terminal and evaluates whether an actual meeting is necessary. Based on this decision, they can hold, cancel, or readjust the agenda for the meeting.

[0139] In this way, we can efficiently resolve the challenges inherent in meetings while also creating a meeting structure that takes into account the emotional state of the users.

[0140] (Example 2)

[0141] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0142] Traditional meeting systems can efficiently gather participants' opinions and perspectives, but they struggle to facilitate discussions while considering the emotional states of individual participants. Furthermore, determining how to utilize meeting results tends to be ambiguous. This leads to problems with the overall quality and efficiency of meetings.

[0143] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0144] In this invention, the server includes a function for inputting information, a function for receiving meeting topics and participants' opinions as data, a function for analyzing the received data and generating a surrogate intelligence agent based on the participants' positions, and a function for measuring the emotional state of the participants and reflecting it in the progress of the discussion. This enables effective meeting management and efficient consensus building that takes into account the emotional information of the participants.

[0145] The "information input function" refers to the part of the system that allows users to input meeting topics and participants' opinions through electronic forms.

[0146] The "function to analyze received data and generate surrogate intelligence agents based on the participants' positions" refers to a function that has a process in which the server analyzes the received data and creates a virtual representative to mimic each participant.

[0147] The "function to conduct meetings in a virtual environment" is a function that allows a generated surrogate intelligence agent to simulate a meeting under computer control.

[0148] The "function to obtain the results of discussions" is a function that collects and records the dialogue and results between agents through virtual meetings.

[0149] The "function to summarize the results of a virtual meeting and output it as summary information" is a function that summarizes the many pieces of information obtained in a virtual meeting, focusing on the most important points, and provides them to the user in a concise manner.

[0150] The "function to measure emotional states and reflect them in the progress of the discussion" refers to a function that analyzes participants' emotions from their voices, facial expressions, etc., and collects information to use in facilitating the meeting based on that analysis.

[0151] A "generative AI model" is a computational model based on artificial intelligence used to generate new knowledge and data.

[0152] A "prompt statement" is a sentence used as input to a generative AI model, serving as an instruction or guideline for the model to respond or generate data.

[0153] This invention is a system aimed at improving the efficiency of meetings and can enhance the quality of discussions by analyzing the emotions of participants. An embodiment of this invention is shown, consisting of a server, terminals, and an emotion engine. This configuration can comprehensively support everything from meeting preparation and execution to results analysis.

[0154] First, users input the meeting agenda and participants' opinions via their devices. A dedicated application is installed on the devices, allowing information to be entered in text or voice format. The entered information is then transmitted to the server using secure communication protocols such as SSL / TLS.

[0155] Next, the server generates an intelligent agent based on the received data. This is where a generative AI model comes into play, utilizing algorithms implemented in Python, Java, etc., to create a surrogate intelligent agent by analyzing past conversation history and received emotional data. Furthermore, by using natural language processing libraries and speech recognition APIs, emotional data is extracted from participants' voices and texts, and this data is analyzed by an emotion engine. This enables the agent to take actions based on emotions.

[0156] As a concrete example, we can consider a product development meeting and use "a situation where the technical department has strong concerns about the market strategy for a new product" as input data. Sentiment analysis based on this data will clarify the technical department's interests and concerns, significantly influencing the progress of the virtual meeting. Examples of prompts could include, "What emotional factors should we be mindful of when discussing product strategy that reflects the technical department's strong concerns?"

[0157] In this way, the server, terminals, and emotion engine work together to provide a system that considers the emotional state of participants and links real-world and virtual meetings, enabling efficient and meaningful meeting management.

[0158] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0159] Step 1:

[0160] The user enters the meeting agenda and participants' opinions into the terminal.

[0161] In terms of specific actions, users use a dedicated app on their device to input information in text or voice format. For example, a user might use the text input function to input "marketing strategy for a new product" as the topic.

[0162] This generates input data on the terminal, which then becomes the target of the next processing step.

[0163] Step 2:

[0164] The terminal sends the data entered by the user to the server.

[0165] The device securely transfers this data to the server using HTTPS.

[0166] Specifically, the terminal converts the input text and audio data into JSON format and sends an HTTP POST request to the server.

[0167] The output of this step is meeting information, which serves as initial data for processing on the server.

[0168] Step 3:

[0169] The server analyzes the received data and generates a surrogate intelligence agent for each participant.

[0170] This process uses a generative AI model that takes into account past conversation history and sentiment data. Specifically, an algorithm implemented in Python handles this, and past meeting data is also retrieved and used from a database.

[0171] The output is an intelligent agent generated based on each participant's position, ready to participate in the virtual meeting in the next step.

[0172] Step 4:

[0173] The server uses an emotion engine to analyze participants' emotions from their voice and text.

[0174] This system utilizes natural language processing libraries and speech recognition APIs to quantify participants' emotional responses. Specifically, the API calculates positive and negative emotion scores from the text.

[0175] The output is emotional information used to facilitate discussions in virtual meetings.

[0176] Step 5:

[0177] The server conducts meetings using intelligent agents in a virtual environment.

[0178] Using a generative AI model, each agent exchanges opinions with others and reaches a conclusion. The agents form opinions that take emotional triggers into consideration.

[0179] The output of this step is the result of the discussion and inferences based on sentiment analysis.

[0180] Step 6:

[0181] The server summarizes the results of the virtual meeting and sends them to the terminal.

[0182] A text generation model is used to create summary information containing key points. Specifically, the server automatically generates the summary using Natural Language Generation (NLG) technology.

[0183] The output is summary information provided to the user.

[0184] Step 7:

[0185] Users review the provided summary information and decide whether or not an actual meeting is necessary.

[0186] Users can use the summary as a reference to schedule the next meeting focusing on a specific issue. Specifically, users review the summary and reschedule the meeting.

[0187] The output of this step is scheduling the meeting and planning supplementary discussions.

[0188] (Application Example 2)

[0189] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0190] In recent years, there has been a growing demand for improved customer experience in physical stores, and personalized customer service that responds to customer emotions has become a crucial factor in enhancing competitiveness. However, under current store operations, it is difficult for sales staff to accurately understand customer emotions in real time and respond appropriately. Therefore, the challenge is to provide a system that accurately recognizes customer emotions and improves the quality of customer service.

[0191] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0192] In this invention, the server includes means for inputting information, means for receiving meeting agendas and participants' opinions as data, means for analyzing the data and generating surrogate intelligence functions appropriate to the participants' positions, and means for acquiring the emotional state of participants in real time using an emotion recognition engine and adjusting the service style and proposals based on that. This enables improved customer satisfaction by responding immediately to customer emotions and providing personalized service.

[0193] "Means of inputting information" refers to the interface through which users provide data such as meeting agendas and participants' opinions to the system.

[0194] "Means of receiving participants' opinions as data" refers to the process by which the system receives opinions and information provided by each individual participating in the meeting.

[0195] "A means of analyzing data and generating surrogate intelligence functions tailored to the participants' perspectives" refers to the process of forming a virtual intelligent agent based on the received data, taking into account each participant's viewpoint and opinions.

[0196] An "emotion recognition engine" is a technology that analyzes participants' emotions from their voice and facial expressions to understand their state in real time.

[0197] "Means for adjusting customer service style and proposals" refers to a function that automatically optimizes responses and product offerings to meet customer needs based on acquired emotional data.

[0198] "Means of providing summaries" refers to the process of providing participants with information that concisely summarizes the results of virtual meetings and sentiment recognition.

[0199] "Means of determining whether an actual meeting is necessary" refers to criteria or processes for evaluating how necessary a physical meeting is, based on the summary information provided.

[0200] To implement this invention, an application system with emotion recognition capabilities is used. The system uses smart glasses to enable interaction with customers. The smart glasses are equipped with a microphone and a camera, which are used to acquire voice and facial expression data in real time.

[0201] The server analyzes this data using an emotion recognition engine to assess the customer's emotions. This process involves using tools such as Microsoft® Azure® Face API to calculate an emotion score from voice tone and facial expressions. Based on the acquired emotion score, the device determines the customer's current emotional state and provides appropriate customer service style suggestions to the salesperson's smart glasses. This allows the salesperson to suggest products that better meet the customer's needs.

[0202] For example, as a customer walks through the store, smart glasses read the customer's facial expressions, and if they detect an emotional score indicating "likely to be interested," they display product information relevant to that customer on the glasses. In this way, the system stimulates customer purchasing intent and maximizes sales opportunities.

[0203] An example of a prompt might be, "Think of ways to analyze a customer's facial expressions and tone of voice to determine their emotional state in real time and adjust the service style accordingly." This prompt is used to demonstrate how the emotion recognition engine should function.

[0204] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0205] Step 1:

[0206] The user puts on smart glasses and begins interacting with the customer. The camera and microphone built into the smart glasses capture the customer's face and voice as input. This collects the data necessary for emotion recognition.

[0207] Step 2:

[0208] The device transmits the acquired audio and video data to the server in real time. The server uses this data as input and processes it with an emotion recognition engine, analyzing the customer's emotions from their voice tone and facial expressions. As a result of the data calculation, an emotion score is generated.

[0209] Step 3:

[0210] The server sends the generated emotion score back to the terminal as output. The terminal receives this emotion score as input and uses it to determine the customer's emotional state. Specifically, a customer service style is selected according to the emotion score.

[0211] Step 4:

[0212] The terminal displays information on the smart glasses' screen based on the selected customer service style. This includes information on products and services recommended to the customer. The terminal adjusts its response in real time based on this information.

[0213] Step 5:

[0214] Based on information from their devices, users can suggest the most suitable products to customers. These suggestions are based on emotional data, enabling sales staff to provide services that better meet customer needs.

[0215] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0216] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0217] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0218] [Second Embodiment]

[0219] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0220] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0221] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0222] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0223] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0224] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0225] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0226] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0227] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0228] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0229] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0230] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0231] In order to implement the present invention, it is necessary to construct a system equipped with fundamental components. The main components of this system include a server, a terminal, and user interfaces. Specific embodiments are described below.

[0232] First, the user inputs the meeting agenda and participants' opinions using a terminal. This data is sent to the server via the terminal. The server receives this data and generates intelligent agents tailored to each participant's position and opinions. The generated AI agents have the ability to mimic each participant's perspective and conduct a virtual meeting.

[0233] The server holds virtual meetings among intelligent agents and gathers information through discussion. The server summarizes the results obtained during this process, creating a human-readable summary. The server then sends this summary to the terminal, which the user can view.

[0234] Based on the provided summary, users can determine whether an actual meeting needs to be held, enabling efficient decision-making. Consider a project management meeting as a concrete example. The user collects opinions from participants in advance and sends them to the server. The server analyzes this data and generates intelligent agents for each participant. In the virtual meeting, the agents exchange opinions and summarize the results. Finally, a statistical analysis report or a summary in presentation format is provided to the user, reducing the need for subsequent discussions.

[0235] This embodiment allows for the smooth preparation and operation of meetings involving large numbers of people, enabling efficient use of work time. The use of this system is also expected to reduce the overall length of meetings.

[0236] The following describes the processing flow.

[0237] Step 1:

[0238] The user inputs the meeting agenda and participants' opinions into a terminal. The terminal converts this into the system's input data format and sends it to the server.

[0239] Step 2:

[0240] The server analyzes the data received from the terminal. Based on the results of the data analysis, it generates intelligent agents tailored to each participant's opinions and perspectives.

[0241] Step 3:

[0242] The server hosts a virtual meeting among the generated intelligent agents. Each agent presents their opinion on the agenda and the discussion progresses through interaction with other agents.

[0243] Step 4:

[0244] The server extracts and summarizes the conclusions and key discussion points obtained during the virtual meeting. The summarized summary is organized to include key points, agreements, and, if necessary, further considerations.

[0245] Step 5:

[0246] The server sends the generated summary to the terminal. The user reviews the summary via the terminal and decides whether an actual meeting needs to be held. If the summary is sufficient, the meeting can be omitted.

[0247] (Example 1)

[0248] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0249] Currently, preparing and running large-scale meetings requires a tremendous amount of time and effort. In particular, gathering participants' opinions in advance and conducting efficient discussions based on those opinions is difficult. Traditional methods often result in insufficient information sharing and consensus building before the meeting, leading to prolonged meetings. Therefore, there is a need for efficient methods of information gathering and discussion beforehand.

[0250] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0251] In this invention, the server includes a device for inputting information, a device means for receiving topics related to the meeting and participants' opinions as digital data, a device means for analyzing the digital data and generating artificial intelligence agents according to the roles of the participants, and a device means for conducting a virtual meeting among the generated artificial intelligence agents and obtaining the results of the discussion. This makes it possible to integrate participants' opinions in advance through the virtual meeting and improve the efficiency of the meeting.

[0252] "A device for inputting information" refers to a device or interface for users to input the topic of a meeting or the opinions of participants.

[0253] A "device for receiving meeting-related topics and participants' opinions as digital data" refers to a device that receives meeting content and participants' opinions in digital format and makes them analyzable.

[0254] A "device that analyzes digital data and generates artificial intelligence agents according to the roles of the participants" refers to a device that creates artificial intelligence agents based on the received digital data, according to the position and opinions of each participant.

[0255] A "device for conducting virtual meetings among generated artificial intelligence agents and obtaining the results of discussions" refers to a device that allows multiple artificial intelligence agents to engage in discussions in a virtual space and obtain the resulting information.

[0256] A "device that provides data as a summary" refers to a device that summarizes and provides the results of a virtual meeting in a format that is easy for humans to understand.

[0257] "A device that determines whether an actual meeting is necessary based on the provided summary data" refers to a device that uses summarized information to determine whether or not a physical meeting needs to be held.

[0258] In order to implement the present invention, it is necessary to construct a system in which each component works in coordination. Specifically, it is important that the server, terminal, and user interface work together.

[0259] Users will input meeting content and participants' opinions using their devices. Personal computers and tablet devices can be used as this interface. The software is expected to include common communication tools such as messaging platforms and document creation tools.

[0260] The terminal sends the entered data to the server. The data is transmitted via a protocol such as HTTPS, ensuring the security of the communication. The data format should preferably be one that can be efficiently processed by the server, such as JSON or XML.

[0261] The server can utilize data processing libraries to analyze the received data. Examples include Python's Pandas and NumPy. Based on the results of this analysis, an artificial intelligence agent is generated using a generative AI model. For the generative AI, a model excelling in natural language processing, such as a large-scale language model, is used.

[0262] The generated artificial intelligence agent mimics the opinions of the participants and conducts a virtual meeting. The server manages this virtual meeting and records the content of the discussion. The information obtained from the discussion is summarized using NLP techniques. The specific techniques used here include OpenAI APIs and other natural language processing frameworks.

[0263] Finally, the generated summary is sent to the device and displayed on the user's device. The user uses this summary as a reference to decide whether or not to hold an actual meeting, if necessary. Throughout this entire process, meeting preparation is made more efficient, and users can make important decisions in a short amount of time.

[0264] As a concrete example, consider a meeting regarding the market launch strategy for a new product. Users collect participants' opinions using their devices, and these opinions are sent to a server. The server builds an AI agent based on each participant's opinion and finds a harmonious solution through a virtual discussion.

[0265] An example of a prompt to input into the generating AI model would be: "Simulate a virtual meeting based on the following meeting agenda and participant opinions, and summarize the results. Agenda: New product launch strategy. Participant opinions: Participant A…, Participant B…"

[0266] Thus, the present invention enables efficient meeting management and allows for the virtual integration of diverse perspectives from participants.

[0267] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0268] Step 1:

[0269] Users input meeting agenda items and participants' opinions into a terminal. Personal computers or tablet devices are used as input interfaces. The input data includes detailed opinions from each participant, thereby collecting the perspectives of the meeting attendees.

[0270] Step 2:

[0271] The terminal converts the input data into JSON format and sends it to the server using the HTTPS protocol. This ensures data security and efficient delivery to the server. The output is data in a format that can be processed by the server.

[0272] Step 3:

[0273] The server parses the received JSON data. This parsing uses data processing libraries such as Python's Pandas and NumPy. Specifically, it organizes and structures each participant's opinion as text data. The output at this stage is a parsed data frame.

[0274] Step 4:

[0275] The server invokes a generative AI model based on the analyzed data to generate artificial intelligence agents that mimic the perspectives of each participant. A large-scale language model is used as the generative AI model. For this generation, prompt sentences reflecting the opinions of each participant are prepared and input into the AI ​​model. The output is multiple artificial intelligence agents.

[0276] Step 5:

[0277] The server initiates a virtual meeting among the generated artificial intelligence agents. Each agent exchanges opinions according to their respective roles and the discussion progresses. Information is collected and stored in real time, and a discussion log is generated as output.

[0278] Step 6:

[0279] The server summarizes the logs obtained from the virtual meeting. Natural language processing (NLP) techniques are used to extract the main points, agreements, and disagreements of the discussion. The output is a human-readable text summary.

[0280] Step 7:

[0281] The server sends the generated summary to the terminal. Again, it is transmitted securely using the HTTPS protocol. The user can receive this summary on their terminal and understand its content. The output is a document or presentation material containing the summary.

[0282] Step 8:

[0283] The user determines whether to actually hold a meeting by using the provided summary. This enables prior integration of opinions, shortening the actual meeting time and making decision-making more efficient. The output is the scheduled meeting based on the user's decision.

[0284] (Application Example 1)

[0285] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0286] In modern factory environments, there is a demand for increased work efficiency and smoother operations. However, communication between multiple devices and participants is required, and there is an issue that it takes time for adjustment. To solve this, it is necessary to generate an optimal work schedule while effectively reflecting the opinions of the participants and support rapid decision-making on-site.

[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0288] In this invention, the server includes a device for inputting information, a device means for receiving the meeting topic and the opinions of the participants as data, a device means for analyzing the data and generating an agent intelligent agent according to the position of the participants, and a device means for inputting on-site environmental data and analyzing the operating status of the device. This enables efficient imitation of the perspectives of each participant and enables rapid and accurate decision-making on-site.

[0289] The "device for inputting information" is a device that enables the user to collect various data and input it into the system.

[0290] The "meeting topic" refers to a specific topic or theme that is the subject of discussion.

[0291] The "opinions of the participants" refer to the proposals and views held individually by each participant in the meeting.

[0292] A "device that receives data" is a device that collects information and data from external sources and incorporates it into a system in an appropriate format.

[0293] An "analytical device" is a device that analyzes input data and extracts the underlying meanings and patterns.

[0294] A "proxy intelligence agent" is a program that aims to mimic the position and opinions of a specific participant and virtually fulfill that role.

[0295] "Virtual dialogue" refers to the exchange of opinions and communication that takes place between agents in a simulated manner.

[0296] A "device for obtaining the results of a discussion" is a device that receives the results of virtual dialogues between agents and aggregates their contents.

[0297] The "device provided as a report" is a device that summarizes the results of a virtual meeting and provides feedback to the user in an easy-to-understand format.

[0298] "On-site environmental data" refers to data that describes the conditions and circumstances at the work site.

[0299] "Equipment operating status" refers to information indicating how production equipment and machinery are currently working.

[0300] An "optimal work schedule" refers to a timetable or plan formulated to efficiently and effectively carry out tasks or projects.

[0301] This invention provides a system for achieving efficient management in a factory. The system aims to digitize on-site work conditions and support rapid decision-making.

[0302] The server is connected to a device for inputting information and receives the theme of the meeting and the opinions of the participants as data. The terminal device (e.g., smart glasses) inputs the environmental data collected from the site in real time. The data input on the terminal is transmitted to the server. The server analyzes the data using Python and generates an agent intelligent agent that mimics the positions of the participants.

[0303] The generated intelligent agent conducts a virtual dialogue based on the generated AI model using the scripted prompt text. The server dynamically obtains the results of this virtual dialogue and summarizes them as a report. The report is provided through the terminal device to prompt appropriate responses at the site.

[0304] As a specific example, when a delay occurs on the production line of a factory, this system analyzes the operating status of each device based on the data collected from the sensors. Then, it generates an optimal work schedule and reports it to the manager. This enables prompt countermeasures.

[0305] An example of the prompt text is "Please detect the current device status of the factory with sensors and create an optimized schedule for the next maintenance time." In this way, the system supports efficient business operations.

[0306] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0307] Step 1:

[0308] The terminal inputs the environmental data and opinions of the site from the user. The input data includes the sensor information of the site and the comments of the manager. This data is transmitted from the terminal to the server.

[0309] Step 2:

[0310] The server analyzes the received data. Using Python, it processes environmental data, including the operating status and signs of anomalies of each device, and extracts patterns and trends. This analysis provides the foundational data for generating intelligent agents.

[0311] Step 3:

[0312] The server generates surrogate intelligent agents using a generative AI model based on preprocessed data. It creates prompt statements and instructs each agent to mimic individual opinions and perspectives. This step takes into account past discussion history and information related to specific topics.

[0313] Step 4:

[0314] The server facilitates virtual dialogue between the generated intelligent agents. Each agent exchanges opinions and advances the discussion based on pre-configured prompts. This process facilitates opinion coordination within the virtual space.

[0315] Step 5:

[0316] The server retrieves the results of the virtual dialogue and summarizes key information and insights. It performs data analysis and extracts key points of agreement and areas requiring further discussion as a summary.

[0317] Step 6:

[0318] The server sends the generated summary to the terminal. The user reviews this summary and makes the best decision based on the situation on site. This final result forms the basis for determining the next action.

[0319] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0320] This invention is a system for improving the efficiency of meetings, and has a configuration that combines a user emotion engine that recognizes user emotions. In this system, the server, terminals, and the emotion analysis engine work together in coordination.

[0321] In terms of usage, users input meeting agenda items and participants' opinions using a terminal. This information is sent to the server, which then generates intelligent agents for each participant based on the received information. During the generation of intelligent agents, the user's emotional data, analyzed by the emotion engine, is also taken into consideration. Specifically, emotional data is obtained from voice and text to understand the emotional responses of the participants.

[0322] The server virtually conducts meetings between intelligent agents, utilizing an emotion engine to gather emotional information and use it to guide the discussion and determine its importance. The intelligent agents can offer opinions and adjust the discussion based on emotional triggers. In this way, the user's emotional state is woven into the flow of the meeting, enabling a more refined and nuanced discussion.

[0323] The results of the virtual meeting are summarized by the server and sent to the terminal as a summary, including emotional feedback. By referring to this summary, users can evaluate the necessity of the actual meeting and make efficient decisions. For example, in a product development meeting, if concerns from the technical department are accompanied by strong emotional reactions, this point will be highlighted in the summary, and the user can hold a separate meeting specifically for that issue.

[0324] This configuration enhances the effectiveness of meetings and provides a new approach for users to advance discussions more quickly and accurately. The feedback system utilizing the emotion engine goes beyond mere opinion gathering and can also take into account the inner state of participants, thus enabling more comprehensive meeting management.

[0325] The following describes the processing flow.

[0326] Step 1:

[0327] The user uses a terminal to input the meeting agenda and the opinions of the participating members. This information may include data in voice or text format. The terminal prepares to send the entered information to the server.

[0328] Step 2:

[0329] The server analyzes the data received from the terminal. This data analysis includes not only the opinions provided by the user, but also the acquisition of various sentiment indicators. The server uses a sentiment engine to recognize the user's emotions from voice and text and add them to the data.

[0330] Step 3:

[0331] The server generates intelligent agents for each participant based on the analysis results. During this generation process, the acquired emotional information is taken into consideration, and each agent is configured to mimic the emotions and perspectives of the participants.

[0332] Step 4:

[0333] The server conducts virtual meetings among the intelligent agents. In these virtual meetings, each agent expresses their opinions and positions and interacts with other agents. Based on emotional information, agents adjust the timing of their counterarguments and endorsements to optimize the flow of the discussion.

[0334] Step 5:

[0335] The server summarizes the results of the virtual meeting. The summarization process extracts key points based on main points of agreement, disagreement, and emotional responses. Emotional feedback is also reflected in the summary, clearly identifying points that require adjustment.

[0336] Step 6:

[0337] The server sends the generated summary to the terminal. The user reviews the summary on the terminal and evaluates whether an actual meeting is necessary. Based on this decision, they can hold, cancel, or readjust the agenda for the meeting.

[0338] In this way, we can efficiently resolve the challenges inherent in meetings while also creating a meeting structure that takes into account the emotional state of the users.

[0339] (Example 2)

[0340] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0341] Traditional meeting systems can efficiently gather participants' opinions and perspectives, but they struggle to facilitate discussions while considering the emotional states of individual participants. Furthermore, determining how to utilize meeting results tends to be ambiguous. This leads to problems with the overall quality and efficiency of meetings.

[0342] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0343] In this invention, the server includes a function for inputting information, a function for receiving meeting topics and participants' opinions as data, a function for analyzing the received data and generating a surrogate intelligence agent based on the participants' positions, and a function for measuring the emotional state of the participants and reflecting it in the progress of the discussion. This enables effective meeting management and efficient consensus building that takes into account the emotional information of the participants.

[0344] The "information input function" refers to the part of the system that allows users to input meeting topics and participants' opinions through electronic forms.

[0345] The "function to analyze received data and generate surrogate intelligence agents based on the participants' positions" refers to a function that has a process in which the server analyzes the received data and creates a virtual representative to mimic each participant.

[0346] The "function to conduct meetings in a virtual environment" is a function that allows a generated surrogate intelligence agent to simulate a meeting under computer control.

[0347] The "function to obtain the results of discussions" is a function that collects and records the dialogue and results between agents through virtual meetings.

[0348] The "function to summarize the results of a virtual meeting and output it as summary information" is a function that summarizes the many pieces of information obtained in a virtual meeting, focusing on the most important points, and provides them to the user in a concise manner.

[0349] The "function to measure emotional states and reflect them in the progress of the discussion" refers to a function that analyzes participants' emotions from their voices, facial expressions, etc., and collects information to use in facilitating the meeting based on that analysis.

[0350] A "generative AI model" is a computational model based on artificial intelligence used to generate new knowledge and data.

[0351] A "prompt statement" is a sentence used as input to a generative AI model, serving as an instruction or guideline for the model to respond or generate data.

[0352] This invention is a system aimed at improving the efficiency of meetings and can enhance the quality of discussions by analyzing the emotions of participants. An embodiment of this invention is shown, consisting of a server, terminals, and an emotion engine. This configuration can comprehensively support everything from meeting preparation and execution to results analysis.

[0353] First, users input the meeting agenda and participants' opinions via their devices. A dedicated application is installed on the devices, allowing information to be entered in text or voice format. The entered information is then transmitted to the server using secure communication protocols such as SSL / TLS.

[0354] Next, the server generates an intelligent agent based on the received data. This is where a generative AI model comes into play, utilizing algorithms implemented in languages ​​such as Python and Java to create a surrogate intelligent agent by analyzing past conversation history and received emotional data. Furthermore, by using natural language processing libraries and speech recognition APIs, emotional data is extracted from participants' voices and texts, and this data is analyzed by an emotion engine. This enables the agent to take actions based on emotions.

[0355] As a concrete example, we can consider a product development meeting and use "a situation where the technical department has strong concerns about the market strategy for a new product" as input data. Sentiment analysis based on this data will clarify the technical department's interests and concerns, significantly influencing the progress of the virtual meeting. Examples of prompts could include, "What emotional factors should we be mindful of when discussing product strategy that reflects the technical department's strong concerns?"

[0356] In this way, the server, terminals, and emotion engine work together to provide a system that considers the emotional state of participants and links real-world and virtual meetings, enabling efficient and meaningful meeting management.

[0357] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0358] Step 1:

[0359] The user enters the meeting agenda and participants' opinions into the terminal.

[0360] In terms of specific actions, users use a dedicated app on their device to input information in text or voice format. For example, a user might use the text input function to input "marketing strategy for a new product" as the topic.

[0361] This generates input data on the terminal, which then becomes the target of the next processing step.

[0362] Step 2:

[0363] The terminal sends the data entered by the user to the server.

[0364] The device securely transfers this data to the server using HTTPS.

[0365] Specifically, the terminal converts the input text and audio data into JSON format and sends an HTTP POST request to the server.

[0366] The output of this step is meeting information, which serves as initial data for processing on the server.

[0367] Step 3:

[0368] The server analyzes the received data and generates a surrogate intelligence agent for each participant.

[0369] This process uses a generative AI model that takes into account past conversation history and sentiment data. Specifically, an algorithm implemented in Python handles this, and past meeting data is also retrieved and used from a database.

[0370] The output is an intelligent agent generated based on each participant's position, ready to participate in the virtual meeting in the next step.

[0371] Step 4:

[0372] The server uses an emotion engine to analyze participants' emotions from their voice and text.

[0373] This system utilizes natural language processing libraries and speech recognition APIs to quantify participants' emotional responses. Specifically, the API calculates positive and negative emotion scores from the text.

[0374] The output is emotional information used to facilitate discussions in virtual meetings.

[0375] Step 5:

[0376] The server conducts meetings using intelligent agents in a virtual environment.

[0377] Using a generative AI model, each agent exchanges opinions with others and reaches a conclusion. The agents form opinions that take emotional triggers into consideration.

[0378] The output of this step is the result of the discussion and inferences based on sentiment analysis.

[0379] Step 6:

[0380] The server summarizes the results of the virtual meeting and sends them to the terminal.

[0381] A text generation model is used to create summary information containing key points. Specifically, the server automatically generates the summary using Natural Language Generation (NLG) technology.

[0382] The output is summary information provided to the user.

[0383] Step 7:

[0384] Users review the provided summary information and decide whether or not an actual meeting is necessary.

[0385] Users can use the summary as a reference to schedule the next meeting focusing on a specific issue. Specifically, users review the summary and reschedule the meeting.

[0386] The output of this step is scheduling the meeting and planning supplementary discussions.

[0387] (Application Example 2)

[0388] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0389] In recent years, there has been a growing demand for improved customer experience in physical stores, and personalized customer service that responds to customer emotions has become a crucial factor in enhancing competitiveness. However, under current store operations, it is difficult for sales staff to accurately understand customer emotions in real time and respond appropriately. Therefore, the challenge is to provide a system that accurately recognizes customer emotions and improves the quality of customer service.

[0390] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0391] In this invention, the server includes means for inputting information, means for receiving meeting agendas and participants' opinions as data, means for analyzing the data and generating surrogate intelligence functions appropriate to the participants' positions, and means for acquiring the emotional state of participants in real time using an emotion recognition engine and adjusting the service style and proposals based on that. This enables improved customer satisfaction by responding immediately to customer emotions and providing personalized service.

[0392] "Means of inputting information" refers to the interface through which users provide data such as meeting agendas and participants' opinions to the system.

[0393] "Means of receiving participants' opinions as data" refers to the process by which the system receives opinions and information provided by each individual participating in the meeting.

[0394] "A means of analyzing data and generating surrogate intelligence functions tailored to the participants' perspectives" refers to the process of forming a virtual intelligent agent based on the received data, taking into account each participant's viewpoint and opinions.

[0395] An "emotion recognition engine" is a technology that analyzes participants' emotions from their voice and facial expressions to understand their state in real time.

[0396] "Means for adjusting customer service style and proposals" refers to a function that automatically optimizes responses and product offerings to meet customer needs based on acquired emotional data.

[0397] "Means of providing summaries" refers to the process of providing participants with information that concisely summarizes the results of virtual meetings and sentiment recognition.

[0398] "Means of determining whether an actual meeting is necessary" refers to criteria or processes for evaluating how necessary a physical meeting is, based on the summary information provided.

[0399] To implement this invention, an application system with emotion recognition capabilities is used. The system uses smart glasses to enable interaction with customers. The smart glasses are equipped with a microphone and a camera, which are used to acquire voice and facial expression data in real time.

[0400] The server analyzes this data using an emotion recognition engine to assess the customer's emotions. This process involves using tools such as Microsoft Azure's Face API to calculate an emotion score from voice tone and facial expressions. Based on the acquired emotion score, the device determines the customer's current emotional state and provides appropriate customer service style suggestions to the salesperson's smart glasses. This allows the salesperson to suggest products that better meet the customer's needs.

[0401] For example, as a customer walks through the store, smart glasses read the customer's facial expressions, and if they detect an emotional score indicating "likely to be interested," they display product information relevant to that customer on the glasses. In this way, the system stimulates customer purchasing intent and maximizes sales opportunities.

[0402] An example of a prompt might be, "Think of ways to analyze a customer's facial expressions and tone of voice to determine their emotional state in real time and adjust the service style accordingly." This prompt is used to demonstrate how the emotion recognition engine should function.

[0403] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0404] Step 1:

[0405] The user puts on smart glasses and begins interacting with the customer. The camera and microphone built into the smart glasses capture the customer's face and voice as input. This collects the data necessary for emotion recognition.

[0406] Step 2:

[0407] The device transmits the acquired audio and video data to the server in real time. The server uses this data as input and processes it with an emotion recognition engine, analyzing the customer's emotions from their voice tone and facial expressions. As a result of the data calculation, an emotion score is generated.

[0408] Step 3:

[0409] The server sends the generated emotion score back to the terminal as output. The terminal receives this emotion score as input and uses it to determine the customer's emotional state. Specifically, a customer service style is selected according to the emotion score.

[0410] Step 4:

[0411] The terminal displays information on the smart glasses' screen based on the selected customer service style. This includes information on products and services recommended to the customer. The terminal adjusts its response in real time based on this information.

[0412] Step 5:

[0413] Based on information from their devices, users can suggest the most suitable products to customers. These suggestions are based on emotional data, enabling sales staff to provide services that better meet customer needs.

[0414] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0415] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0416] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0417] [Third Embodiment]

[0418] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0419] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0420] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0421] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0422] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0423] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0424] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0425] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0426] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0427] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0428] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0429] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0430] In order to implement the present invention, it is necessary to construct a system equipped with fundamental components. The main components of this system include a server, a terminal, and user interfaces. Specific embodiments are described below.

[0431] First, the user inputs the meeting agenda and participants' opinions using a terminal. This data is sent to the server via the terminal. The server receives this data and generates intelligent agents tailored to each participant's position and opinions. The generated AI agents have the ability to mimic each participant's perspective and conduct a virtual meeting.

[0432] The server holds virtual meetings among intelligent agents and gathers information through discussion. The server summarizes the results obtained during this process, creating a human-readable summary. The server then sends this summary to the terminal, which the user can view.

[0433] Based on the provided summary, users can determine whether an actual meeting needs to be held, enabling efficient decision-making. Consider a project management meeting as a concrete example. The user collects opinions from participants in advance and sends them to the server. The server analyzes this data and generates intelligent agents for each participant. In the virtual meeting, the agents exchange opinions and summarize the results. Finally, a statistical analysis report or a summary in presentation format is provided to the user, reducing the need for subsequent discussions.

[0434] This embodiment allows for the smooth preparation and operation of meetings involving large numbers of people, enabling efficient use of work time. The use of this system is also expected to reduce the overall length of meetings.

[0435] The following describes the processing flow.

[0436] Step 1:

[0437] The user inputs the meeting agenda and participants' opinions into a terminal. The terminal converts this into the system's input data format and sends it to the server.

[0438] Step 2:

[0439] The server analyzes the data received from the terminal. Based on the results of the data analysis, it generates intelligent agents tailored to each participant's opinions and perspectives.

[0440] Step 3:

[0441] The server hosts a virtual meeting among the generated intelligent agents. Each agent presents their opinion on the agenda and the discussion progresses through interaction with other agents.

[0442] Step 4:

[0443] The server extracts and summarizes the conclusions and key discussion points obtained during the virtual meeting. The summarized summary is organized to include key points, agreements, and, if necessary, further considerations.

[0444] Step 5:

[0445] The server sends the generated summary to the terminal. The user reviews the summary via the terminal and decides whether an actual meeting needs to be held. If the summary is sufficient, the meeting can be omitted.

[0446] (Example 1)

[0447] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0448] Currently, preparing and running large-scale meetings requires a tremendous amount of time and effort. In particular, gathering participants' opinions in advance and conducting efficient discussions based on those opinions is difficult. Traditional methods often result in insufficient information sharing and consensus building before the meeting, leading to prolonged meetings. Therefore, there is a need for efficient methods of information gathering and discussion beforehand.

[0449] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0450] In this invention, the server includes a device for inputting information, a device means for receiving topics related to the meeting and participants' opinions as digital data, a device means for analyzing the digital data and generating artificial intelligence agents according to the roles of the participants, and a device means for conducting a virtual meeting among the generated artificial intelligence agents and obtaining the results of the discussion. This makes it possible to integrate participants' opinions in advance through the virtual meeting and improve the efficiency of the meeting.

[0451] "A device for inputting information" refers to a device or interface for users to input the topic of a meeting or the opinions of participants.

[0452] A "device for receiving meeting-related topics and participants' opinions as digital data" refers to a device that receives meeting content and participants' opinions in digital format and makes them analyzable.

[0453] A "device that analyzes digital data and generates artificial intelligence agents according to the roles of the participants" refers to a device that creates artificial intelligence agents based on the received digital data, according to the position and opinions of each participant.

[0454] A "device for conducting virtual meetings among generated artificial intelligence agents and obtaining the results of discussions" refers to a device that allows multiple artificial intelligence agents to engage in discussions in a virtual space and obtain the resulting information.

[0455] A "device that provides data as a summary" refers to a device that summarizes and provides the results of a virtual meeting in a format that is easy for humans to understand.

[0456] "A device that determines whether an actual meeting is necessary based on the provided summary data" refers to a device that uses summarized information to determine whether or not a physical meeting needs to be held.

[0457] In order to implement the present invention, it is necessary to construct a system in which each component works in coordination. Specifically, it is important that the server, terminal, and user interface work together.

[0458] Users will input meeting content and participants' opinions using their devices. Personal computers and tablet devices can be used as this interface. The software is expected to include common communication tools such as messaging platforms and document creation tools.

[0459] The terminal sends the entered data to the server. The data is transmitted via a protocol such as HTTPS, ensuring the security of the communication. The data format should preferably be one that can be efficiently processed by the server, such as JSON or XML.

[0460] The server can utilize data processing libraries to analyze the received data. Examples include Python's Pandas and NumPy. Based on the results of this analysis, an artificial intelligence agent is generated using a generative AI model. For the generative AI, a model excelling in natural language processing, such as a large-scale language model, is used.

[0461] The generated artificial intelligence agent mimics the opinions of the participants and conducts a virtual meeting. The server manages this virtual meeting and records the content of the discussion. The information obtained from the discussion is summarized using NLP techniques. The specific techniques used here include OpenAI APIs and other natural language processing frameworks.

[0462] Finally, the generated summary is sent to the device and displayed on the user's device. The user uses this summary as a reference to decide whether or not to hold an actual meeting, if necessary. Throughout this entire process, meeting preparation is made more efficient, and users can make important decisions in a short amount of time.

[0463] As a concrete example, consider a meeting regarding the market launch strategy for a new product. Users collect participants' opinions using their devices, and these opinions are sent to a server. The server builds an AI agent based on each participant's opinion and finds a harmonious solution through a virtual discussion.

[0464] An example of a prompt to input into the generating AI model would be: "Simulate a virtual meeting based on the following meeting agenda and participant opinions, and summarize the results. Agenda: New product launch strategy. Participant opinions: Participant A…, Participant B…"

[0465] Thus, the present invention enables efficient meeting management and allows for the virtual integration of diverse perspectives from participants.

[0466] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0467] Step 1:

[0468] Users input meeting agenda items and participants' opinions into a terminal. Personal computers or tablet devices are used as input interfaces. The input data includes detailed opinions from each participant, thereby collecting the perspectives of the meeting attendees.

[0469] Step 2:

[0470] The terminal converts the input data into JSON format and sends it to the server using the HTTPS protocol. This ensures data security and efficient delivery to the server. The output is data in a format that can be processed by the server.

[0471] Step 3:

[0472] The server parses the received JSON data. This parsing uses data processing libraries such as Python's Pandas and NumPy. Specifically, it organizes and structures each participant's opinion as text data. The output at this stage is a parsed data frame.

[0473] Step 4:

[0474] The server invokes a generative AI model based on the analyzed data to generate artificial intelligence agents that mimic the perspectives of each participant. A large-scale language model is used as the generative AI model. For this generation, prompt sentences reflecting the opinions of each participant are prepared and input into the AI ​​model. The output is multiple artificial intelligence agents.

[0475] Step 5:

[0476] The server initiates a virtual meeting among the generated artificial intelligence agents. Each agent exchanges opinions according to their respective roles and the discussion progresses. Information is collected and stored in real time, and a discussion log is generated as output.

[0477] Step 6:

[0478] The server summarizes the logs obtained from the virtual meeting. Natural language processing (NLP) techniques are used to extract the main points, agreements, and disagreements of the discussion. The output is a human-readable text summary.

[0479] Step 7:

[0480] The server sends the generated summary to the terminal. Again, it is transmitted securely using the HTTPS protocol. The user can receive this summary on their terminal and understand its content. The output is a document or presentation material containing the summary.

[0481] Step 8:

[0482] Users use the provided summary to decide whether or not to actually hold a meeting. This allows for prior consensus building, reduces actual meeting time, and makes decision-making more efficient. The output is a meeting schedule based on the user's decision.

[0483] (Application Example 1)

[0484] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0485] Modern factory environments demand increased work efficiency and smoother operations. However, this presents challenges due to the need for communication between multiple pieces of equipment and participants, which can be time-consuming to coordinate. To address this, it is necessary to effectively incorporate participants' opinions, generate optimal work schedules, and support rapid decision-making on-site.

[0486] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0487] In this invention, the server includes a device for inputting information, a device for receiving the meeting topic and participants' opinions as data, a device for analyzing the data and generating a surrogate intelligent agent appropriate to the participants' positions, and a device for inputting on-site environmental data and analyzing the operating status of the device. This enables efficient imitation of each participant's perspective and allows for rapid and accurate decision-making on-site.

[0488] A "device for inputting information" is a device that allows users to collect various types of data and input them into a system.

[0489] The "topic of the meeting" refers to the specific topic or theme that will be discussed.

[0490] "Participant opinions" refer to the individual suggestions and views that each participant holds during the meeting.

[0491] A "device that receives data" is a device that collects information and data from external sources and incorporates it into a system in an appropriate format.

[0492] An "analytical device" is a device that analyzes input data and extracts the underlying meanings and patterns.

[0493] A "proxy intelligence agent" is a program that aims to mimic the position and opinions of a specific participant and virtually fulfill that role.

[0494] "Virtual dialogue" refers to the exchange of opinions and communication that takes place between agents in a simulated manner.

[0495] A "device for obtaining the results of a discussion" is a device that receives the results of virtual dialogues between agents and aggregates their contents.

[0496] The "device provided as a report" is a device that summarizes the results of a virtual meeting and provides feedback to the user in an easy-to-understand format.

[0497] "On-site environmental data" refers to data that describes the conditions and circumstances at the work site.

[0498] "Equipment operating status" refers to information indicating how production equipment and machinery are currently working.

[0499] An "optimal work schedule" refers to a timetable or plan formulated to efficiently and effectively carry out tasks or projects.

[0500] This invention provides a system for achieving efficient management in a factory. The system aims to digitize on-site work conditions and support rapid decision-making.

[0501] The server is connected to an information input device and receives meeting topics and participants' opinions as data. Terminal devices (e.g., smart glasses) input environmental data collected from the field in real time. The data entered on the terminal is sent to the server. The server uses Python to analyze the data and generates a surrogate intelligent agent that mimics the perspective of the participants.

[0502] The generated intelligent agent uses embellished prompts to conduct a virtual dialogue based on the generated AI model. The server dynamically retrieves the results of this virtual dialogue and summarizes them as a report. The report is provided through a terminal device to facilitate appropriate responses on-site.

[0503] As a concrete example, if a delay occurs on a factory production line, this system analyzes the operating status of each piece of equipment based on data collected from sensors. It then generates an optimal work schedule and reports it to the manager. This enables swift countermeasures.

[0504] An example of a prompt message is, "Use sensors to detect the current status of the factory equipment and create a schedule to optimize the next maintenance." In this way, the system supports efficient business operations.

[0505] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0506] Step 1:

[0507] The terminal receives on-site environmental data and comments from the user. The entered data includes on-site sensor information and comments from the administrator. This data is then sent from the terminal to the server.

[0508] Step 2:

[0509] The server analyzes the received data. Using Python, it processes environmental data, including the operating status and signs of anomalies of each device, and extracts patterns and trends. This analysis provides the foundational data for generating intelligent agents.

[0510] Step 3:

[0511] The server generates surrogate intelligent agents using a generative AI model based on preprocessed data. It creates prompt statements and instructs each agent to mimic individual opinions and perspectives. This step takes into account past discussion history and information related to specific topics.

[0512] Step 4:

[0513] The server facilitates virtual dialogue between the generated intelligent agents. Each agent exchanges opinions and advances the discussion based on pre-configured prompts. This process facilitates opinion coordination within the virtual space.

[0514] Step 5:

[0515] The server retrieves the results of the virtual dialogue and summarizes key information and insights. It performs data analysis and extracts key points of agreement and areas requiring further discussion as a summary.

[0516] Step 6:

[0517] The server sends the generated summary to the terminal. The user reviews this summary and makes the best decision based on the situation on site. This final result forms the basis for determining the next action.

[0518] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0519] This invention is a system for improving the efficiency of meetings, and has a configuration that combines a user emotion engine that recognizes user emotions. In this system, the server, terminals, and the emotion analysis engine work together in coordination.

[0520] In terms of usage, users input meeting agenda items and participants' opinions using a terminal. This information is sent to the server, which then generates intelligent agents for each participant based on the received information. During the generation of intelligent agents, the user's emotional data, analyzed by the emotion engine, is also taken into consideration. Specifically, emotional data is obtained from voice and text to understand the emotional responses of the participants.

[0521] The server virtually conducts meetings between intelligent agents, utilizing an emotion engine to gather emotional information and use it to guide the discussion and determine its importance. The intelligent agents can offer opinions and adjust the discussion based on emotional triggers. In this way, the user's emotional state is woven into the flow of the meeting, enabling a more refined and nuanced discussion.

[0522] The results of the virtual meeting are summarized by the server and sent to the terminal as a summary, including emotional feedback. By referring to this summary, users can evaluate the necessity of the actual meeting and make efficient decisions. For example, in a product development meeting, if concerns from the technical department are accompanied by strong emotional reactions, this point will be highlighted in the summary, and the user can hold a separate meeting specifically for that issue.

[0523] This configuration enhances the effectiveness of meetings and provides a new approach for users to advance discussions more quickly and accurately. The feedback system utilizing the emotion engine goes beyond mere opinion gathering and can also take into account the inner state of participants, thus enabling more comprehensive meeting management.

[0524] The following describes the processing flow.

[0525] Step 1:

[0526] The user uses a terminal to input the meeting agenda and the opinions of the participating members. This information may include data in voice or text format. The terminal prepares to send the entered information to the server.

[0527] Step 2:

[0528] The server analyzes the data received from the terminal. This data analysis includes not only the opinions provided by the user, but also the acquisition of various sentiment indicators. The server uses a sentiment engine to recognize the user's emotions from voice and text and add them to the data.

[0529] Step 3:

[0530] The server generates intelligent agents for each participant based on the analysis results. During this generation process, the acquired emotional information is taken into consideration, and each agent is configured to mimic the emotions and perspectives of the participants.

[0531] Step 4:

[0532] The server conducts virtual meetings among the intelligent agents. In these virtual meetings, each agent expresses their opinions and positions and interacts with other agents. Based on emotional information, agents adjust the timing of their counterarguments and endorsements to optimize the flow of the discussion.

[0533] Step 5:

[0534] The server summarizes the results of the virtual meeting. The summarization process extracts key points based on main points of agreement, disagreement, and emotional responses. Emotional feedback is also reflected in the summary, clearly identifying points that require adjustment.

[0535] Step 6:

[0536] The server sends the generated summary to the terminal. The user reviews the summary on the terminal and evaluates whether an actual meeting is necessary. Based on this decision, they can hold, cancel, or readjust the agenda for the meeting.

[0537] In this way, we can efficiently resolve the challenges inherent in meetings while also creating a meeting structure that takes into account the emotional state of the users.

[0538] (Example 2)

[0539] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0540] Traditional meeting systems can efficiently gather participants' opinions and perspectives, but they struggle to facilitate discussions while considering the emotional states of individual participants. Furthermore, determining how to utilize meeting results tends to be ambiguous. This leads to problems with the overall quality and efficiency of meetings.

[0541] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0542] In this invention, the server includes a function for inputting information, a function for receiving meeting topics and participants' opinions as data, a function for analyzing the received data and generating a surrogate intelligence agent based on the participants' positions, and a function for measuring the emotional state of the participants and reflecting it in the progress of the discussion. This enables effective meeting management and efficient consensus building that takes into account the emotional information of the participants.

[0543] The "information input function" refers to the part of the system that allows users to input meeting topics and participants' opinions through electronic forms.

[0544] The "function to analyze received data and generate surrogate intelligence agents based on the participants' positions" refers to a function that has a process in which the server analyzes the received data and creates a virtual representative to mimic each participant.

[0545] The "function to conduct meetings in a virtual environment" is a function that allows a generated surrogate intelligence agent to simulate a meeting under computer control.

[0546] The "function to obtain the results of discussions" is a function that collects and records the dialogue and results between agents through virtual meetings.

[0547] The "function to summarize the results of a virtual meeting and output it as summary information" is a function that summarizes the many pieces of information obtained in a virtual meeting, focusing on the most important points, and provides them to the user in a concise manner.

[0548] The "function to measure emotional states and reflect them in the progress of the discussion" refers to a function that analyzes participants' emotions from their voices, facial expressions, etc., and collects information to use in facilitating the meeting based on that analysis.

[0549] A "generative AI model" is a computational model based on artificial intelligence used to generate new knowledge and data.

[0550] A "prompt statement" is a sentence used as input to a generative AI model, serving as an instruction or guideline for the model to respond or generate data.

[0551] This invention is a system aimed at improving the efficiency of meetings and can enhance the quality of discussions by analyzing the emotions of participants. An embodiment of this invention is shown, consisting of a server, terminals, and an emotion engine. This configuration can comprehensively support everything from meeting preparation and execution to results analysis.

[0552] First, users input the meeting agenda and participants' opinions via their devices. A dedicated application is installed on the devices, allowing information to be entered in text or voice format. The entered information is then transmitted to the server using secure communication protocols such as SSL / TLS.

[0553] Next, the server generates an intelligent agent based on the received data. This is where a generative AI model comes into play, utilizing algorithms implemented in languages ​​such as Python and Java to create a surrogate intelligent agent by analyzing past conversation history and received emotional data. Furthermore, by using natural language processing libraries and speech recognition APIs, emotional data is extracted from participants' voices and texts, and this data is analyzed by an emotion engine. This enables the agent to take actions based on emotions.

[0554] As a concrete example, we can consider a product development meeting and use "a situation where the technical department has strong concerns about the market strategy for a new product" as input data. Sentiment analysis based on this data will clarify the technical department's interests and concerns, significantly influencing the progress of the virtual meeting. Examples of prompts could include, "What emotional factors should we be mindful of when discussing product strategy that reflects the technical department's strong concerns?"

[0555] In this way, the server, terminals, and emotion engine work together to provide a system that considers the emotional state of participants and links real-world and virtual meetings, enabling efficient and meaningful meeting management.

[0556] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0557] Step 1:

[0558] The user enters the meeting agenda and participants' opinions into the terminal.

[0559] In terms of specific actions, users use a dedicated app on their device to input information in text or voice format. For example, a user might use the text input function to input "marketing strategy for a new product" as the topic.

[0560] This generates input data on the terminal, which then becomes the target of the next processing step.

[0561] Step 2:

[0562] The terminal sends the data entered by the user to the server.

[0563] The device securely transfers this data to the server using HTTPS.

[0564] Specifically, the terminal converts the input text and audio data into JSON format and sends an HTTP POST request to the server.

[0565] The output of this step is meeting information, which serves as initial data for processing on the server.

[0566] Step 3:

[0567] The server analyzes the received data and generates a surrogate intelligence agent for each participant.

[0568] This process uses a generative AI model that takes into account past conversation history and sentiment data. Specifically, an algorithm implemented in Python handles this, and past meeting data is also retrieved and used from a database.

[0569] The output is an intelligent agent generated based on each participant's position, ready to participate in the virtual meeting in the next step.

[0570] Step 4:

[0571] The server uses an emotion engine to analyze participants' emotions from their voice and text.

[0572] This system utilizes natural language processing libraries and speech recognition APIs to quantify participants' emotional responses. Specifically, the API calculates positive and negative emotion scores from the text.

[0573] The output is emotional information used to facilitate discussions in virtual meetings.

[0574] Step 5:

[0575] The server conducts meetings using intelligent agents in a virtual environment.

[0576] Using a generative AI model, each agent exchanges opinions with others and reaches a conclusion. The agents form opinions that take emotional triggers into consideration.

[0577] The output of this step is the result of the discussion and inferences based on sentiment analysis.

[0578] Step 6:

[0579] The server summarizes the results of the virtual meeting and sends them to the terminal.

[0580] A text generation model is used to create summary information containing key points. Specifically, the server automatically generates the summary using Natural Language Generation (NLG) technology.

[0581] The output is summary information provided to the user.

[0582] Step 7:

[0583] Users review the provided summary information and decide whether or not an actual meeting is necessary.

[0584] Users can use the summary as a reference to schedule the next meeting focusing on a specific issue. Specifically, users review the summary and reschedule the meeting.

[0585] The output of this step is scheduling the meeting and planning supplementary discussions.

[0586] (Application Example 2)

[0587] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0588] In recent years, there has been a growing demand for improved customer experience in physical stores, and personalized customer service that responds to customer emotions has become a crucial factor in enhancing competitiveness. However, under current store operations, it is difficult for sales staff to accurately understand customer emotions in real time and respond appropriately. Therefore, the challenge is to provide a system that accurately recognizes customer emotions and improves the quality of customer service.

[0589] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0590] In this invention, the server includes means for inputting information, means for receiving meeting agendas and participants' opinions as data, means for analyzing the data and generating surrogate intelligence functions appropriate to the participants' positions, and means for acquiring the emotional state of participants in real time using an emotion recognition engine and adjusting the service style and proposals based on that. This enables improved customer satisfaction by responding immediately to customer emotions and providing personalized service.

[0591] "Means of inputting information" refers to the interface through which users provide data such as meeting agendas and participants' opinions to the system.

[0592] "Means of receiving participants' opinions as data" refers to the process by which the system receives opinions and information provided by each individual participating in the meeting.

[0593] "A means of analyzing data and generating surrogate intelligence functions tailored to the participants' perspectives" refers to the process of forming a virtual intelligent agent based on the received data, taking into account each participant's viewpoint and opinions.

[0594] An "emotion recognition engine" is a technology that analyzes participants' emotions from their voice and facial expressions to understand their state in real time.

[0595] "Means for adjusting customer service style and proposals" refers to a function that automatically optimizes responses and product offerings to meet customer needs based on acquired emotional data.

[0596] "Means of providing summaries" refers to the process of providing participants with information that concisely summarizes the results of virtual meetings and sentiment recognition.

[0597] "Means of determining whether an actual meeting is necessary" refers to criteria or processes for evaluating how necessary a physical meeting is, based on the summary information provided.

[0598] To implement this invention, an application system with emotion recognition capabilities is used. The system uses smart glasses to enable interaction with customers. The smart glasses are equipped with a microphone and a camera, which are used to acquire voice and facial expression data in real time.

[0599] The server analyzes this data using an emotion recognition engine to assess the customer's emotions. This process involves using tools such as Microsoft Azure's Face API to calculate an emotion score from voice tone and facial expressions. Based on the acquired emotion score, the device determines the customer's current emotional state and provides appropriate customer service style suggestions to the salesperson's smart glasses. This allows the salesperson to suggest products that better meet the customer's needs.

[0600] For example, as a customer walks through the store, smart glasses read the customer's facial expressions, and if they detect an emotional score indicating "likely to be interested," they display product information relevant to that customer on the glasses. In this way, the system stimulates customer purchasing intent and maximizes sales opportunities.

[0601] An example of a prompt might be, "Think of ways to analyze a customer's facial expressions and tone of voice to determine their emotional state in real time and adjust the service style accordingly." This prompt is used to demonstrate how the emotion recognition engine should function.

[0602] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0603] Step 1:

[0604] The user puts on smart glasses and begins interacting with the customer. The camera and microphone built into the smart glasses capture the customer's face and voice as input. This collects the data necessary for emotion recognition.

[0605] Step 2:

[0606] The device transmits the acquired audio and video data to the server in real time. The server uses this data as input and processes it with an emotion recognition engine, analyzing the customer's emotions from their voice tone and facial expressions. As a result of the data calculation, an emotion score is generated.

[0607] Step 3:

[0608] The server sends the generated emotion score back to the terminal as output. The terminal receives this emotion score as input and uses it to determine the customer's emotional state. Specifically, a customer service style is selected according to the emotion score.

[0609] Step 4:

[0610] The terminal displays information on the smart glasses' screen based on the selected customer service style. This includes information on products and services recommended to the customer. The terminal adjusts its response in real time based on this information.

[0611] Step 5:

[0612] Based on information from their devices, users can suggest the most suitable products to customers. These suggestions are based on emotional data, enabling sales staff to provide services that better meet customer needs.

[0613] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0614] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0615] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0616] [Fourth Embodiment]

[0617] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0618] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0619] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0620] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0621] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0622] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0623] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0624] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0625] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0626] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0627] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0628] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0629] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0630] In order to implement the present invention, it is necessary to construct a system equipped with fundamental components. The main components of this system include a server, a terminal, and user interfaces. Specific embodiments are described below.

[0631] First, the user inputs the meeting agenda and participants' opinions using a terminal. This data is sent to the server via the terminal. The server receives this data and generates intelligent agents tailored to each participant's position and opinions. The generated AI agents have the ability to mimic each participant's perspective and conduct a virtual meeting.

[0632] The server holds virtual meetings among intelligent agents and gathers information through discussion. The server summarizes the results obtained during this process, creating a human-readable summary. The server then sends this summary to the terminal, which the user can view.

[0633] Based on the provided summary, users can determine whether an actual meeting needs to be held, enabling efficient decision-making. Consider a project management meeting as a concrete example. The user collects opinions from participants in advance and sends them to the server. The server analyzes this data and generates intelligent agents for each participant. In the virtual meeting, the agents exchange opinions and summarize the results. Finally, a statistical analysis report or a summary in presentation format is provided to the user, reducing the need for subsequent discussions.

[0634] This embodiment allows for the smooth preparation and operation of meetings involving large numbers of people, enabling efficient use of work time. The use of this system is also expected to reduce the overall length of meetings.

[0635] The following describes the processing flow.

[0636] Step 1:

[0637] The user inputs the meeting agenda and participants' opinions into a terminal. The terminal converts this into the system's input data format and sends it to the server.

[0638] Step 2:

[0639] The server analyzes the data received from the terminal. Based on the results of the data analysis, it generates intelligent agents tailored to each participant's opinions and perspectives.

[0640] Step 3:

[0641] The server hosts a virtual meeting among the generated intelligent agents. Each agent presents their opinion on the agenda and the discussion progresses through interaction with other agents.

[0642] Step 4:

[0643] The server extracts and summarizes the conclusions and key discussion points obtained during the virtual meeting. The summarized summary is organized to include key points, agreements, and, if necessary, further considerations.

[0644] Step 5:

[0645] The server sends the generated summary to the terminal. The user reviews the summary via the terminal and decides whether an actual meeting needs to be held. If the summary is sufficient, the meeting can be omitted.

[0646] (Example 1)

[0647] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0648] Currently, preparing and running large-scale meetings requires a tremendous amount of time and effort. In particular, gathering participants' opinions in advance and conducting efficient discussions based on those opinions is difficult. Traditional methods often result in insufficient information sharing and consensus building before the meeting, leading to prolonged meetings. Therefore, there is a need for efficient methods of information gathering and discussion beforehand.

[0649] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0650] In this invention, the server includes a device for inputting information, a device means for receiving topics related to the meeting and participants' opinions as digital data, a device means for analyzing the digital data and generating artificial intelligence agents according to the roles of the participants, and a device means for conducting a virtual meeting among the generated artificial intelligence agents and obtaining the results of the discussion. This makes it possible to integrate participants' opinions in advance through the virtual meeting and improve the efficiency of the meeting.

[0651] "A device for inputting information" refers to a device or interface for users to input the topic of a meeting or the opinions of participants.

[0652] A "device for receiving meeting-related topics and participants' opinions as digital data" refers to a device that receives meeting content and participants' opinions in digital format and makes them analyzable.

[0653] A "device that analyzes digital data and generates artificial intelligence agents according to the roles of the participants" refers to a device that creates artificial intelligence agents based on the received digital data, according to the position and opinions of each participant.

[0654] A "device for conducting virtual meetings among generated artificial intelligence agents and obtaining the results of discussions" refers to a device that allows multiple artificial intelligence agents to engage in discussions in a virtual space and obtain the resulting information.

[0655] A "device that provides data as a summary" refers to a device that summarizes and provides the results of a virtual meeting in a format that is easy for humans to understand.

[0656] "A device that determines whether an actual meeting is necessary based on the provided summary data" refers to a device that uses summarized information to determine whether or not a physical meeting needs to be held.

[0657] In order to implement the present invention, it is necessary to construct a system in which each component works in coordination. Specifically, it is important that the server, terminal, and user interface work together.

[0658] Users will input meeting content and participants' opinions using their devices. Personal computers and tablet devices can be used as this interface. The software is expected to include common communication tools such as messaging platforms and document creation tools.

[0659] The terminal sends the entered data to the server. The data is transmitted via a protocol such as HTTPS, ensuring the security of the communication. The data format should preferably be one that can be efficiently processed by the server, such as JSON or XML.

[0660] The server can utilize data processing libraries to analyze the received data. Examples include Python's Pandas and NumPy. Based on the results of this analysis, an artificial intelligence agent is generated using a generative AI model. For the generative AI, a model excelling in natural language processing, such as a large-scale language model, is used.

[0661] The generated artificial intelligence agent mimics the opinions of the participants and conducts a virtual meeting. The server manages this virtual meeting and records the content of the discussion. The information obtained from the discussion is summarized using NLP techniques. The specific techniques used here include OpenAI APIs and other natural language processing frameworks.

[0662] Finally, the generated summary is sent to the device and displayed on the user's device. The user uses this summary as a reference to decide whether or not to hold an actual meeting, if necessary. Throughout this entire process, meeting preparation is made more efficient, and users can make important decisions in a short amount of time.

[0663] As a concrete example, consider a meeting regarding the market launch strategy for a new product. Users collect participants' opinions using their devices, and these opinions are sent to a server. The server builds an AI agent based on each participant's opinion and finds a harmonious solution through a virtual discussion.

[0664] An example of a prompt to input into the generating AI model would be: "Simulate a virtual meeting based on the following meeting agenda and participant opinions, and summarize the results. Agenda: New product launch strategy. Participant opinions: Participant A…, Participant B…"

[0665] Thus, the present invention enables efficient meeting management and allows for the virtual integration of diverse perspectives from participants.

[0666] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0667] Step 1:

[0668] Users input meeting agenda items and participants' opinions into a terminal. Personal computers or tablet devices are used as input interfaces. The input data includes detailed opinions from each participant, thereby collecting the perspectives of the meeting attendees.

[0669] Step 2:

[0670] The terminal converts the input data into JSON format and sends it to the server using the HTTPS protocol. This ensures data security and efficient delivery to the server. The output is data in a format that can be processed by the server.

[0671] Step 3:

[0672] The server parses the received JSON data. This parsing uses data processing libraries such as Python's Pandas and NumPy. Specifically, it organizes and structures each participant's opinion as text data. The output at this stage is a parsed data frame.

[0673] Step 4:

[0674] The server invokes a generative AI model based on the analyzed data to generate artificial intelligence agents that mimic the perspectives of each participant. A large-scale language model is used as the generative AI model. For this generation, prompt sentences reflecting the opinions of each participant are prepared and input into the AI ​​model. The output is multiple artificial intelligence agents.

[0675] Step 5:

[0676] The server initiates a virtual meeting among the generated artificial intelligence agents. Each agent exchanges opinions according to their respective roles and the discussion progresses. Information is collected and stored in real time, and a discussion log is generated as output.

[0677] Step 6:

[0678] The server summarizes the logs obtained from the virtual meeting. Natural language processing (NLP) techniques are used to extract the main points, agreements, and disagreements of the discussion. The output is a human-readable text summary.

[0679] Step 7:

[0680] The server sends the generated summary to the terminal. Again, it is transmitted securely using the HTTPS protocol. The user can receive this summary on their terminal and understand its content. The output is a document or presentation material containing the summary.

[0681] Step 8:

[0682] Users use the provided summary to decide whether or not to actually hold a meeting. This allows for prior consensus building, reduces actual meeting time, and makes decision-making more efficient. The output is a meeting schedule based on the user's decision.

[0683] (Application Example 1)

[0684] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0685] Modern factory environments demand increased work efficiency and smoother operations. However, this presents challenges due to the need for communication between multiple pieces of equipment and participants, which can be time-consuming to coordinate. To address this, it is necessary to effectively incorporate participants' opinions, generate optimal work schedules, and support rapid decision-making on-site.

[0686] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0687] In this invention, the server includes a device for inputting information, a device for receiving the meeting topic and participants' opinions as data, a device for analyzing the data and generating a surrogate intelligent agent appropriate to the participants' positions, and a device for inputting on-site environmental data and analyzing the operating status of the device. This enables efficient imitation of each participant's perspective and allows for rapid and accurate decision-making on-site.

[0688] A "device for inputting information" is a device that allows users to collect various types of data and input them into a system.

[0689] The "topic of the meeting" refers to the specific topic or theme that will be discussed.

[0690] "Participant opinions" refer to the individual suggestions and views that each participant holds during the meeting.

[0691] A "device that receives data" is a device that collects information and data from external sources and incorporates it into a system in an appropriate format.

[0692] An "analytical device" is a device that analyzes input data and extracts the underlying meanings and patterns.

[0693] A "proxy intelligence agent" is a program that aims to mimic the position and opinions of a specific participant and virtually fulfill that role.

[0694] "Virtual dialogue" refers to the exchange of opinions and communication that takes place between agents in a simulated manner.

[0695] A "device for obtaining the results of a discussion" is a device that receives the results of virtual dialogues between agents and aggregates their contents.

[0696] The "device provided as a report" is a device that summarizes the results of a virtual meeting and provides feedback to the user in an easy-to-understand format.

[0697] "On-site environmental data" refers to data that describes the conditions and circumstances at the work site.

[0698] "Equipment operating status" refers to information indicating how production equipment and machinery are currently working.

[0699] An "optimal work schedule" refers to a timetable or plan formulated to efficiently and effectively carry out tasks or projects.

[0700] This invention provides a system for achieving efficient management in a factory. The system aims to digitize on-site work conditions and support rapid decision-making.

[0701] The server is connected to an information input device and receives meeting topics and participants' opinions as data. Terminal devices (e.g., smart glasses) input environmental data collected from the field in real time. The data entered on the terminal is sent to the server. The server uses Python to analyze the data and generates a surrogate intelligent agent that mimics the perspective of the participants.

[0702] The generated intelligent agent uses embellished prompts to conduct a virtual dialogue based on the generated AI model. The server dynamically retrieves the results of this virtual dialogue and summarizes them as a report. The report is provided through a terminal device to facilitate appropriate responses on-site.

[0703] As a concrete example, if a delay occurs on a factory production line, this system analyzes the operating status of each piece of equipment based on data collected from sensors. It then generates an optimal work schedule and reports it to the manager. This enables swift countermeasures.

[0704] An example of a prompt message is, "Use sensors to detect the current status of the factory equipment and create a schedule to optimize the next maintenance." In this way, the system supports efficient business operations.

[0705] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0706] Step 1:

[0707] The terminal receives on-site environmental data and comments from the user. The entered data includes on-site sensor information and comments from the administrator. This data is then sent from the terminal to the server.

[0708] Step 2:

[0709] The server analyzes the received data. Using Python, it processes environmental data, including the operating status and signs of anomalies of each device, and extracts patterns and trends. This analysis provides the foundational data for generating intelligent agents.

[0710] Step 3:

[0711] The server generates surrogate intelligent agents using a generative AI model based on preprocessed data. It creates prompt statements and instructs each agent to mimic individual opinions and perspectives. This step takes into account past discussion history and information related to specific topics.

[0712] Step 4:

[0713] The server facilitates virtual dialogue between the generated intelligent agents. Each agent exchanges opinions and advances the discussion based on pre-configured prompts. This process facilitates opinion coordination within the virtual space.

[0714] Step 5:

[0715] The server retrieves the results of the virtual dialogue and summarizes key information and insights. It performs data analysis and extracts key points of agreement and areas requiring further discussion as a summary.

[0716] Step 6:

[0717] The server sends the generated summary to the terminal. The user reviews this summary and makes the best decision based on the situation on site. This final result forms the basis for determining the next action.

[0718] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0719] This invention is a system for improving the efficiency of meetings, and has a configuration that combines a user emotion engine that recognizes user emotions. In this system, the server, terminals, and the emotion analysis engine work together in coordination.

[0720] In terms of usage, users input meeting agenda items and participants' opinions using a terminal. This information is sent to the server, which then generates intelligent agents for each participant based on the received information. During the generation of intelligent agents, the user's emotional data, analyzed by the emotion engine, is also taken into consideration. Specifically, emotional data is obtained from voice and text to understand the emotional responses of the participants.

[0721] The server virtually conducts meetings between intelligent agents, utilizing an emotion engine to gather emotional information and use it to guide the discussion and determine its importance. The intelligent agents can offer opinions and adjust the discussion based on emotional triggers. In this way, the user's emotional state is woven into the flow of the meeting, enabling a more refined and nuanced discussion.

[0722] The results of the virtual meeting are summarized by the server and sent to the terminal as a summary, including emotional feedback. By referring to this summary, users can evaluate the necessity of the actual meeting and make efficient decisions. For example, in a product development meeting, if concerns from the technical department are accompanied by strong emotional reactions, this point will be highlighted in the summary, and the user can hold a separate meeting specifically for that issue.

[0723] This configuration enhances the effectiveness of meetings and provides a new approach for users to advance discussions more quickly and accurately. The feedback system utilizing the emotion engine goes beyond mere opinion gathering and can also take into account the inner state of participants, thus enabling more comprehensive meeting management.

[0724] The following describes the processing flow.

[0725] Step 1:

[0726] The user uses a terminal to input the meeting agenda and the opinions of the participating members. This information may include data in voice or text format. The terminal prepares to send the entered information to the server.

[0727] Step 2:

[0728] The server analyzes the data received from the terminal. This data analysis includes not only the opinions provided by the user, but also the acquisition of various sentiment indicators. The server uses a sentiment engine to recognize the user's emotions from voice and text and add them to the data.

[0729] Step 3:

[0730] The server generates intelligent agents for each participant based on the analysis results. During this generation process, the acquired emotional information is taken into consideration, and each agent is configured to mimic the emotions and perspectives of the participants.

[0731] Step 4:

[0732] The server conducts virtual meetings among the intelligent agents. In these virtual meetings, each agent expresses their opinions and positions and interacts with other agents. Based on emotional information, agents adjust the timing of their counterarguments and endorsements to optimize the flow of the discussion.

[0733] Step 5:

[0734] The server summarizes the results of the virtual meeting. The summarization process extracts key points based on main points of agreement, disagreement, and emotional responses. Emotional feedback is also reflected in the summary, clearly identifying points that require adjustment.

[0735] Step 6:

[0736] The server sends the generated summary to the terminal. The user reviews the summary on the terminal and evaluates whether an actual meeting is necessary. Based on this decision, they can hold, cancel, or readjust the agenda for the meeting.

[0737] In this way, we can efficiently resolve the challenges inherent in meetings while also creating a meeting structure that takes into account the emotional state of the users.

[0738] (Example 2)

[0739] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0740] Traditional meeting systems can efficiently gather participants' opinions and perspectives, but they struggle to facilitate discussions while considering the emotional states of individual participants. Furthermore, determining how to utilize meeting results tends to be ambiguous. This leads to problems with the overall quality and efficiency of meetings.

[0741] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0742] In this invention, the server includes a function for inputting information, a function for receiving meeting topics and participants' opinions as data, a function for analyzing the received data and generating a surrogate intelligence agent based on the participants' positions, and a function for measuring the emotional state of the participants and reflecting it in the progress of the discussion. This enables effective meeting management and efficient consensus building that takes into account the emotional information of the participants.

[0743] The "information input function" refers to the part of the system that allows users to input meeting topics and participants' opinions through electronic forms.

[0744] The "function to analyze received data and generate surrogate intelligence agents based on the participants' positions" refers to a function that has a process in which the server analyzes the received data and creates a virtual representative to mimic each participant.

[0745] The "function to conduct meetings in a virtual environment" is a function that allows a generated surrogate intelligence agent to simulate a meeting under computer control.

[0746] The "function to obtain the results of discussions" is a function that collects and records the dialogue and results between agents through virtual meetings.

[0747] The "function to summarize the results of a virtual meeting and output it as summary information" is a function that summarizes the many pieces of information obtained in a virtual meeting, focusing on the most important points, and provides them to the user in a concise manner.

[0748] The "function to measure emotional states and reflect them in the progress of the discussion" refers to a function that analyzes participants' emotions from their voices, facial expressions, etc., and collects information to use in facilitating the meeting based on that analysis.

[0749] A "generative AI model" is a computational model based on artificial intelligence used to generate new knowledge and data.

[0750] A "prompt statement" is a sentence used as input to a generative AI model, serving as an instruction or guideline for the model to respond or generate data.

[0751] This invention is a system aimed at improving the efficiency of meetings and can enhance the quality of discussions by analyzing the emotions of participants. An embodiment of this invention is shown, consisting of a server, terminals, and an emotion engine. This configuration can comprehensively support everything from meeting preparation and execution to results analysis.

[0752] First, users input the meeting agenda and participants' opinions via their devices. A dedicated application is installed on the devices, allowing information to be entered in text or voice format. The entered information is then transmitted to the server using secure communication protocols such as SSL / TLS.

[0753] Next, the server generates an intelligent agent based on the received data. This is where a generative AI model comes into play, utilizing algorithms implemented in languages ​​such as Python and Java to create a surrogate intelligent agent by analyzing past conversation history and received emotional data. Furthermore, by using natural language processing libraries and speech recognition APIs, emotional data is extracted from participants' voices and texts, and this data is analyzed by an emotion engine. This enables the agent to take actions based on emotions.

[0754] As a concrete example, we can consider a product development meeting and use "a situation where the technical department has strong concerns about the market strategy for a new product" as input data. Sentiment analysis based on this data will clarify the technical department's interests and concerns, significantly influencing the progress of the virtual meeting. Examples of prompts could include, "What emotional factors should we be mindful of when discussing product strategy that reflects the technical department's strong concerns?"

[0755] In this way, the server, terminals, and emotion engine work together to provide a system that considers the emotional state of participants and links real-world and virtual meetings, enabling efficient and meaningful meeting management.

[0756] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0757] Step 1:

[0758] The user enters the meeting agenda and participants' opinions into the terminal.

[0759] In terms of specific actions, users use a dedicated app on their device to input information in text or voice format. For example, a user might use the text input function to input "marketing strategy for a new product" as the topic.

[0760] This generates input data on the terminal, which then becomes the target of the next processing step.

[0761] Step 2:

[0762] The terminal sends the data entered by the user to the server.

[0763] The device securely transfers this data to the server using HTTPS.

[0764] Specifically, the terminal converts the input text and audio data into JSON format and sends an HTTP POST request to the server.

[0765] The output of this step is meeting information, which serves as initial data for processing on the server.

[0766] Step 3:

[0767] The server analyzes the received data and generates a surrogate intelligence agent for each participant.

[0768] This process uses a generative AI model that takes into account past conversation history and sentiment data. Specifically, an algorithm implemented in Python handles this, and past meeting data is also retrieved and used from a database.

[0769] The output is an intelligent agent generated based on each participant's position, ready to participate in the virtual meeting in the next step.

[0770] Step 4:

[0771] The server uses an emotion engine to analyze participants' emotions from their voice and text.

[0772] This system utilizes natural language processing libraries and speech recognition APIs to quantify participants' emotional responses. Specifically, the API calculates positive and negative emotion scores from the text.

[0773] The output is emotional information used to facilitate discussions in virtual meetings.

[0774] Step 5:

[0775] The server conducts meetings using intelligent agents in a virtual environment.

[0776] Using a generative AI model, each agent exchanges opinions with others and reaches a conclusion. The agents form opinions that take emotional triggers into consideration.

[0777] The output of this step is the result of the discussion and inferences based on sentiment analysis.

[0778] Step 6:

[0779] The server summarizes the results of the virtual meeting and sends them to the terminal.

[0780] A text generation model is used to create summary information containing key points. Specifically, the server automatically generates the summary using Natural Language Generation (NLG) technology.

[0781] The output is summary information provided to the user.

[0782] Step 7:

[0783] Users review the provided summary information and decide whether or not an actual meeting is necessary.

[0784] Users can use the summary as a reference to schedule the next meeting focusing on a specific issue. Specifically, users review the summary and reschedule the meeting.

[0785] The output of this step is scheduling the meeting and planning supplementary discussions.

[0786] (Application Example 2)

[0787] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0788] In recent years, there has been a growing demand for improved customer experience in physical stores, and personalized customer service that responds to customer emotions has become a crucial factor in enhancing competitiveness. However, under current store operations, it is difficult for sales staff to accurately understand customer emotions in real time and respond appropriately. Therefore, the challenge is to provide a system that accurately recognizes customer emotions and improves the quality of customer service.

[0789] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0790] In this invention, the server includes means for inputting information, means for receiving meeting agendas and participants' opinions as data, means for analyzing the data and generating surrogate intelligence functions appropriate to the participants' positions, and means for acquiring the emotional state of participants in real time using an emotion recognition engine and adjusting the service style and proposals based on that. This enables improved customer satisfaction by responding immediately to customer emotions and providing personalized service.

[0791] "Means of inputting information" refers to the interface through which users provide data such as meeting agendas and participants' opinions to the system.

[0792] "Means of receiving participants' opinions as data" refers to the process by which the system receives opinions and information provided by each individual participating in the meeting.

[0793] "A means of analyzing data and generating surrogate intelligence functions tailored to the participants' perspectives" refers to the process of forming a virtual intelligent agent based on the received data, taking into account each participant's viewpoint and opinions.

[0794] An "emotion recognition engine" is a technology that analyzes participants' emotions from their voice and facial expressions to understand their state in real time.

[0795] "Means for adjusting customer service style and proposals" refers to a function that automatically optimizes responses and product offerings to meet customer needs based on acquired emotional data.

[0796] "Means of providing summaries" refers to the process of providing participants with information that concisely summarizes the results of virtual meetings and sentiment recognition.

[0797] "Means of determining whether an actual meeting is necessary" refers to criteria or processes for evaluating how necessary a physical meeting is, based on the summary information provided.

[0798] To implement this invention, an application system with emotion recognition capabilities is used. The system uses smart glasses to enable interaction with customers. The smart glasses are equipped with a microphone and a camera, which are used to acquire voice and facial expression data in real time.

[0799] The server analyzes this data using an emotion recognition engine to assess the customer's emotions. This process involves using tools such as Microsoft Azure's Face API to calculate an emotion score from voice tone and facial expressions. Based on the acquired emotion score, the device determines the customer's current emotional state and provides appropriate customer service style suggestions to the salesperson's smart glasses. This allows the salesperson to suggest products that better meet the customer's needs.

[0800] For example, as a customer walks through the store, smart glasses read the customer's facial expressions, and if they detect an emotional score indicating "likely to be interested," they display product information relevant to that customer on the glasses. In this way, the system stimulates customer purchasing intent and maximizes sales opportunities.

[0801] An example of a prompt might be, "Think of ways to analyze a customer's facial expressions and tone of voice to determine their emotional state in real time and adjust the service style accordingly." This prompt is used to demonstrate how the emotion recognition engine should function.

[0802] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0803] Step 1:

[0804] The user puts on smart glasses and begins interacting with the customer. The camera and microphone built into the smart glasses capture the customer's face and voice as input. This collects the data necessary for emotion recognition.

[0805] Step 2:

[0806] The device transmits the acquired audio and video data to the server in real time. The server uses this data as input and processes it with an emotion recognition engine, analyzing the customer's emotions from their voice tone and facial expressions. As a result of the data calculation, an emotion score is generated.

[0807] Step 3:

[0808] The server sends the generated emotion score back to the terminal as output. The terminal receives this emotion score as input and uses it to determine the customer's emotional state. Specifically, a customer service style is selected according to the emotion score.

[0809] Step 4:

[0810] The terminal displays information on the smart glasses' screen based on the selected customer service style. This includes information on products and services recommended to the customer. The terminal adjusts its response in real time based on this information.

[0811] Step 5:

[0812] Based on information from their devices, users can suggest the most suitable products to customers. These suggestions are based on emotional data, enabling sales staff to provide services that better meet customer needs.

[0813] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0814] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0815] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0816] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0817] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0818] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0819] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0820] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0821] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0822] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0823] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0824] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0825] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0826] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0827] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0828] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0829] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0830] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0831] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0832] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0833] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0834] The following is further disclosed regarding the embodiments described above.

[0835] (Claim 1)

[0836] A means of inputting information, and a means of receiving the meeting agenda and participants' opinions as data,

[0837] A means for analyzing the aforementioned data and generating a surrogate intelligent agent appropriate to the participant's position,

[0838] A means of conducting a virtual meeting among the generated intelligent agents and obtaining the results of the discussion,

[0839] A means of summarizing the results of a virtual meeting and providing them as a summary,

[0840] Based on the summary provided, a means to determine whether an actual meeting is necessary,

[0841] A system that includes this.

[0842] (Claim 2)

[0843] The system according to claim 1, wherein the intelligent agent has means of using a generative model to consider past conversation history and to mimic the perspective of each participant.

[0844] (Claim 3)

[0845] The system according to claim 1, wherein the summary includes key points of agreement, differences, and points requiring further discussion.

[0846] "Example 1"

[0847] (Claim 1)

[0848] A device for inputting information, and a device for receiving the topics related to the meeting and the opinions of the participants as digital data,

[0849] The apparatus includes means for analyzing the aforementioned digital data and generating artificial intelligence agents according to the roles of the participants,

[0850] A device and means for conducting virtual meetings among generated artificial intelligence agents and obtaining the results of discussions,

[0851] A device and means for summarizing the results of a virtual meeting and providing them as summary data,

[0852] A device that determines whether an actual meeting is necessary based on the provided summary data,

[0853] A system that includes this.

[0854] (Claim 2)

[0855] The system according to claim 1, wherein the artificial intelligence agent has a device that uses generation technology to consider past conversation history and reproduce the perspective of each participant.

[0856] (Claim 3)

[0857] The apparatus system according to claim 1, which includes in the summary data key points of agreement, differences, and points requiring further discussion.

[0858] "Application Example 1"

[0859] (Claim 1)

[0860] A device for inputting information, and a device for receiving the meeting topic and participants' opinions as data,

[0861] The apparatus includes means for analyzing the aforementioned data and generating a surrogate intelligent agent appropriate to the participant's position,

[0862] A device means for conducting virtual dialogue between generated intelligent agents and obtaining the results of the discussion,

[0863] A device and means for summarizing the results of a virtual meeting and providing them as a report,

[0864] A device and means for determining whether an actual meeting is necessary based on the provided report,

[0865] A device that inputs on-site environmental data and analyzes the operating status of the equipment,

[0866] A device and means for generating an optimal work schedule based on operating conditions,

[0867] A system that includes this.

[0868] (Claim 2)

[0869] The system according to claim 1, wherein the intelligent agent has a device that uses a generative model to consider past conversation history and mimics the perspective of each participant.

[0870] (Claim 3)

[0871] The apparatus system according to claim 1, which includes in the aforementioned report the main points of agreement, differences, and points requiring further discussion.

[0872] "Example 2 of combining an emotion engine"

[0873] (Claim 1)

[0874] A function for inputting information, and a means for receiving meeting topics and participants' opinions as data,

[0875] A functional means for analyzing received data and generating a surrogate intelligence agent based on the participant's position,

[0876] A functional means for conducting meetings in a virtual environment among generated intelligent agents and obtaining the results of the discussions,

[0877] A functional means for summarizing the results of a virtual meeting and outputting them as summary information,

[0878] A functional means for determining whether an actual meeting is necessary based on the summary information provided,

[0879] A functional means to measure the emotional state of participants and reflect it in the progress of the discussion,

[0880] A system that includes this.

[0881] (Claim 2)

[0882] The system according to claim 1, wherein the intelligent agent has the function of mimicking the perspective of each participant by utilizing a generative AI model and taking into account past conversation history and sentiment data.

[0883] (Claim 3)

[0884] The system according to claim 1, which includes in summary information key points of agreement, differences, and points requiring further consideration.

[0885] "Application example 2 when combining with an emotional engine"

[0886] (Claim 1)

[0887] A means of inputting information, and a means of receiving the meeting agenda and participants' opinions as data,

[0888] A means for analyzing the aforementioned data and generating surrogate intelligence functions according to the participant's position,

[0889] A means of conducting a virtual meeting among the generated intelligent functions and obtaining the results of the discussion,

[0890] A means of acquiring participants' emotional states in real time using an emotion recognition engine and adjusting customer service style and proposals based on that,

[0891] A means of summarizing the results of a virtual meeting and providing them as a summary,

[0892] Based on the summary provided, a means to determine whether an actual meeting is necessary,

[0893] A system that includes this.

[0894] (Claim 2)

[0895] The intelligent function has means of using a generative model to consider past conversation history and to mimic the perspective of each participant, as described in claim 1.

[0896] (Claim 3)

[0897] The system according to claim 1, wherein the summary includes key points of agreement, differences, and points requiring further discussion, as well as a method for adjusting the automated customer service style. [Explanation of Symbols]

[0898] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of inputting information, and a means of receiving the meeting agenda and participants' opinions as data, A means for analyzing the aforementioned data and generating a surrogate intelligent agent appropriate to the participant's position, A means of conducting a virtual meeting among the generated intelligent agents and obtaining the results of the discussion, A means of summarizing the results of a virtual meeting and providing them as a summary, Based on the summary provided, a means to determine whether an actual meeting is necessary, A system that includes this.

2. The system according to claim 1, wherein the intelligent agent has means for using a generative model to consider past conversation history and to mimic the perspective of each participant.

3. The system according to claim 1, wherein the summary includes key points of agreement, differences, and points requiring further discussion.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A