System

The system addresses inconsistent sales negotiation quality and inefficiencies by converting voice input to text, using generative AI for summaries and proposals, and integrating progress management and information search, resulting in standardized negotiations and efficient information retrieval.

JP2026015081APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116555
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing sales negotiation systems lack consistency in quality, require significant time and effort for progress management and information search, leading to inefficiencies.

Method used

A system that converts user voice or text input into real-time text data, uses generative AI for summary and proposal generation, and integrates progress management and information search functionalities.

Benefits of technology

Standardizes business negotiation quality, enhances progress management efficiency, and enables quick and accurate information acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026015081000001_ABST
    Figure 2026015081000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for a user to provide instructions through audio input and for a device to convert the audio to text in real time; means for sending the textual AI to a server and for the server to generate summaries of information and suggestions using production-based data; and means for the device to display the summaries and suggestions sent from the server to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Describe the "problem that the invention aims to solve" and the "means for solving the problem."

[0005] Sales professionals are required to respond efficiently and consistently to a wide range of tasks, including sales negotiations, progress management, and information search. However, with current methods, the quality of sales negotiations tends to vary from person to person, and progress management and information search often require time and effort. For this reason, a method is needed to achieve both efficiency in sales activities and uniform quality. [Means for solving the problem]

[0006] The present invention standardizes the quality of business negotiations by providing a system that includes a means for a user to give instructions through voice input and for a terminal to convert the voice into text in real time, a means for transmitting the text data to a server, and for the server to generate a summary and proposal of the information using generative AI, and a means for the terminal to display the summary and proposal transmitted from the server to the user.

[0007] Furthermore, the efficiency of progress management is improved by a system including a means for the user to input progress by voice or text and the terminal to send it to a server, a means for the server to update the progress management database and generate the next action or alert, and a means for the terminal to notify the user of the generated action or alert.

[0008] In addition, the system includes a means for a user to search for specific information by voice or text and for the terminal to send that information to a server, a means for the server to search for related information from a database, generate search results and send them to the terminal, and a means for the terminal to display the search results to the user, enabling quick and accurate information acquisition.

[0009] Understood. Below are definitions for each of the key terms contained in the claims.

[0010] A "user" is a person who operates the system and gives instructions by voice or text.

[0011] "Voice input" is a means of recognizing and processing the user's voice as digital data.

[0012] A "terminal" is a device operated by a user that converts voice input into text, communicates with a server, and displays the results.

[0013] "Speech recognition" is a technology in which a terminal receives a user's voice and converts it into text data.

[0014] "Text data" refers to character information generated by speech recognition and sent to the server.

[0015] The "server" is a remote computer system that uses generative AI to analyze text data and generate information summaries and suggestions.

[0016] "Generative AI" is an artificial intelligence technology that processes received data and generates useful information and suggestions for users.

[0017] "Information summary" is data that generative AI extracts important points from text data and summarizes them in a concise form.

[0018] A "proposal" is a presentation of the next action or solution that the server creates using generative AI based on past data and patterns.

[0019] "Progress" refers to the progress and degree of achievement of business activities.

[0020] The "progress management database" is a database in which the server records and manages progress information.

[0021] "Action" is a suggestion that indicates the specific operation or behavior that the user should take next depending on the progress.

[0022] An "alert" is a system feature that notifies the user when a specific condition is not met or of important notifications.

[0023] "Information search" is an operation in which a user inputs a keyword to search for specific information and obtains related data.

[0024] A "database" is a collection of information that a server stores and allows for searching and retrieval as needed.

[0025] "Search results" are relevant information that the server extracts from its database based on the user's query. [Brief explanation of the drawings]

[0026] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0027] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0028] First, the terms used in the following description will be explained.

[0029] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0030] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0031] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0032] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0033] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0034] [First embodiment]

[0035] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0036] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0037] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0038] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0039] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0040] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0041] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0042] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0043] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0044] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0045] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0046] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0047] Understood. Below is the "Form for carrying out the invention".

[0048] In one embodiment of the invention, the system begins with a user providing instructions through voice input, which the device converts in real time into text and sends to the server. The server then uses generative AI to analyze the text data, generating summaries and suggestions, which are then sent back to the device for display to the user. When the user inputs progress, this is recorded in a progress management database, and the user is notified of next actions or alerts. Furthermore, when the user searches for specific information, the server searches for related data and provides the results to the user.

[0049] Business negotiation support

[0050] When a user launches the app, the app activates its voice input function. The device converts the user's speech into text in real time and sends the text data to a server. The server then uses generative AI to analyze the text data, summarize key points, and suggest next steps. These summaries and suggestions are displayed to the user on the device.

[0051] For example, suppose a user is introducing a new product during a business meeting and the client asks, "What is the price of this product?" In this case, the device recognizes the voice as a "price question" and sends the text data to the server. The server analyzes past data, generates appropriate price information, and returns it to the device. As a result, the user can respond to the client by saying, "The price of this product is XX yen."

[0052] Progress management support

[0053] In progress management, the user inputs the progress of each business negotiation or task by voice or text. The device recognizes this and sends the data to the server. The server updates the progress management database and generates the next action or alert. These actions and alerts are notified to the user via the device.

[0054] For example, when a user says, "Set up a meeting with Client A next week," the device converts the voice to text and sends it to the server. The server updates the progress management database and generates an alert to set up the next meeting. The device notifies the user when the meeting date approaches.

[0055] Information search support

[0056] With the information search function, users search for the information they need using voice or text. The device analyzes this input and sends the query to the server. The server then searches for relevant information in a database, generates results, and sends them to the device. The user can then view the search results through the device.

[0057] For example, if a user says, "I want to know the technical specifications of a new product," the device converts the instruction into text and sends it to the server. The server then searches the database for the new product's technical specifications and returns them to the device. The user can then view the specifications on the screen.

[0058] As a result, the system standardizes the quality of business negotiations, improves the efficiency of progress management, and enables quick and accurate information acquisition.

[0059] The processing flow will be explained below.

[0060] Understood. Below are the specific processing steps.

[0061] Business negotiation support

[0062] Step 1:

[0063] The user launches the app by tapping the app icon.

[0064] Step 2:

[0065] The terminal activates the voice input function and starts recording the user's voice in real time.

[0066] Step 3:

[0067] The device converts the recorded audio into text data in real time, using a speech recognition algorithm to output the audio as text.

[0068] Step 4:

[0069] The terminal transmits the text data to the server, and the text data is uploaded to the server via the network.

[0070] Step 5:

[0071] The server receives the text data and uses generative AI to summarize the information, including important points and next actions.

[0072] Step 6:

[0073] The server generates information summaries and suggestions, and creates optimal suggestions based on the results of analyzing the text data.

[0074] Step 7:

[0075] The server transmits the generated summary data and the proposal to the terminal, and transmits the data to the terminal via the network.

[0076] Step 8:

[0077] The device displays summary data and suggestions to the user, either as text on the screen or as audio feedback.

[0078] Progress management support

[0079] Step 1:

[0080] The user types a progress update into the app or gives a voice command. They tap the progress update button and enter the information.

[0081] Step 2:

[0082] The terminal performs voice recognition or text analysis on the input content to generate text data.

[0083] Step 3:

[0084] The device sends the text data to the server, where it is uploaded via the network.

[0085] Step 4:

[0086] The server receives the progress data and updates the progress management database, recording the new progress information.

[0087] Step 5:

[0088] The server analyzes the progress and generates the next action or alert: Create the necessary tasks or alerts.

[0089] Step 6:

[0090] The server sends the generated tasks and alerts to the terminal and returns the data via the network.

[0091] Step 7:

[0092] The device notifies the user of tasks and alerts, displaying them on the screen and providing audio feedback where appropriate.

[0093] Information search support

[0094] Step 1:

[0095] When a user wants to search for specific information, they input it by voice or text, then tap the search button.

[0096] Step 2:

[0097] The terminal performs voice recognition or text analysis on the input content to generate text data.

[0098] Step 3:

[0099] The device sends the text data to the server, where it is uploaded via the network.

[0100] Step 4:

[0101] The server receives the query and searches the database to extract information that matches the specified keywords.

[0102] Step 5:

[0103] The server generates search results and sends them to the device, organizes the results, and sends the data to the device.

[0104] Step 6:

[0105] The device displays the search results to the user, displaying them on the screen and providing audio feedback if necessary.

[0106] By following the above steps, this system provides functions such as support for business negotiations, progress management, and information search.

[0107] Example 1

[0108] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0109] Conventional sales negotiation support systems, progress management systems, and information search systems each function independently, resulting in low user convenience. Furthermore, there is a lack of systems that can efficiently convert voice input into text and then centrally manage the subsequent analysis, proposals, progress management, and information search. This creates challenges for users, making it difficult to standardize the quality of sales negotiations, efficiently manage progress, and quickly obtain information.

[0110] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0111] In this invention, the server includes means for analyzing information using a generative artificial intelligence model and generating summaries and proposals, means for updating the progress management database and generating next actions and alerts, and means for searching the database for related information, generating search results, and sending them to the terminal. This allows for centralized management of business negotiation support, progress management, and information search, enabling users to efficiently give instructions and obtain information.

[0112] "User" refers to any individual or corporation that uses this system.

[0113] "Voice input" refers to a means by which a user gives instructions to a system using voice.

[0114] "Terminal" refers to a hardware device that converts speech to text, transmits the data to a server, and displays the results.

[0115] "Server" refers to a computer system that uses generative artificial intelligence models to analyze text data and generate summaries, suggestions, progress management, and information search results.

[0116] A "generative artificial intelligence model" refers to an algorithm or program that analyzes input text data and generates summaries or suggestions based on it.

[0117] "Text data" refers to data resulting from converting voice input into text.

[0118] "Analyzing information" refers to the process of understanding input text data and generating summaries or suggestions based on that content.

[0119] A "summary" refers to a concise summary of the important points of text data.

[0120] "Suggestion" refers to showing the user the next action or countermeasure to be taken based on the analysis results.

[0121] A "progress management database" refers to a database for recording and managing the progress of business negotiations and tasks.

[0122] "Action" refers to the specific next steps or actions to be taken.

[0123] "Alert" refers to a warning or reminder that notifies the user of important information or progress.

[0124] "Information search" refers to a search performed by a user to obtain specific information.

[0125] "Database" refers to a data storage system in which related information is stored.

[0126] "Search Results" refers to a server-generated answer or collection of data in response to an information search query.

[0127] This invention is a system that unifies the management of business negotiation support, progress management, and information search. This system is composed of three main components: users, terminals, and servers.

[0128] Business negotiation support

[0129] When a user launches the app, the device activates the voice input function. When the user speaks, the device converts the speech into text in real time and sends the text data to the server. The server then analyzes the text data using a generative artificial intelligence model (e.g., OpenAI's GPT series), summarizes the key points, and suggests the next action to take. The suggested summary and action are displayed to the user on the device.

[0130] For example, if a user is introducing a new product and the client asks, "What is the price of this product?", the device converts the speech into text "Question about price" and sends it to the server. The server analyzes past data, generates appropriate price information, and returns it to the device. The user can then reply to the client, "The price of this product is XX yen."

[0131] Progress management

[0132] When the user inputs progress by voice or text, the device sends the data to the server, which updates the progress management database (e.g., PostgreSQL) and generates the next action or alert. The generated action or alert is then notified to the user via the device.

[0133] For example, if a user says, "Set up a meeting with Client A next week," the device converts the speech into text and sends it to the server. The server updates the progress management database and generates an alert to set up the next meeting. When the meeting date approaches, the device notifies the user.

[0134] Information Search

[0135] When a user searches for the information they need by voice or text, the device sends the query to the server, which then searches for relevant information in a database (e.g., Elasticsearch), generates results, and sends them to the device. The user can then view the search results through the device.

[0136] For example, if a user says, "I want to know the technical specifications of a new product," the device converts the instruction into text and sends it to the server. The server then searches the database for the new product's technical specifications and returns them to the device. The user can then view the specifications on the device's screen.

[0137] As a result, this system can standardize the quality of business negotiations, improve the efficiency of progress management, and enable quick and accurate information acquisition.

[0138] Prompt Sentence Examples

[0139] For sales support: "Please summarize the notes from your meeting with Client X and tell me what action to take next."

[0140] For progress management support: "How is project Y progressing? What's the next action?"

[0141] For information search assistance: "Please tell me the technical specifications for new product Z."

[0142] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0143] Understood. Below we will explain in detail the processing flow of the system program, divided into processing steps.

[0144] Business negotiation support

[0145] Step 1:

[0146] A user launches an app and activates the voice input function. The input is the command to launch the app, and the output is that voice input is enabled.

[0147] Step 2:

[0148] The user provides voice input, which the device converts into text in real time. The input is the user's voice data, and the output is text data. Specifically, the device uses voice recognition software (e.g., Google Speech-to-Text API).

[0149] Step 3:

[0150] The terminal sends the generated text data to the server. The input is the text data, and the output is the data sent to the server. Specifically, the terminal sends data to the server using an HTTPS request.

[0151] Step 4:

[0152] The server uses a generative AI model to analyze the transmitted text data. The input is text data, and the output is a summary and proposed data. Specifically, the server uses a generative AI model such as GPT-4 to perform the analysis and generation process.

[0153] Step 5:

[0154] The server sends the generated summary and proposal to the terminal. The input is the summary and proposal data, and the output is the data to be sent to the terminal. Specifically, the server sends the generated results in JSON format.

[0155] Step 6:

[0156] The terminal displays the received summary and suggestions to the user. The input is the data received from the server, and the output is the visualization to the user. Specifically, the terminal displays the data using a GUI.

[0157] Progress management

[0158] Step 1:

[0159] The user inputs their progress by voice or text. The input is the user's voice or text data, and the output is text data. Specifically, the device recognizes the voice and converts it into text.

[0160] Step 2:

[0161] The terminal sends the text data to the server. The input is the text data, and the output is the data sent to the server. Specific operations use an HTTPS request.

[0162] Step 3:

[0163] The server updates the progress management database. The input is text data, and the output is the updated database. Specifically, the server updates the database using an SQL query.

[0164] Step 4:

[0165] The server generates the next action or alert. The input is the updated database information, and the output is the action / alert data. Specifically, the server checks for pending tasks and generates the next action or alert.

[0166] Step 5:

[0167] The device notifies the user of the generated actions and alerts. The input is the action / alert data sent from the server, and the output is the notification to the user. Specifically, the device performs push notifications and in-app notifications.

[0168] Information Search

[0169] Step 1:

[0170] The user searches for the information they need by voice or text. The input is the user's query (voice or text), and the output is the query data. Specifically, the device performs voice recognition and converts it into text.

[0171] Step 2:

[0172] The device sends the query to the server. The input is the query data, and the output is the data sent to the server. Specific operations use an HTTPS request.

[0173] Step 3:

[0174] The server searches for relevant information from a database. The input is the query data, and the output is the search result data. Specifically, the server uses a search engine such as Elasticsearch.

[0175] Step 4:

[0176] The server generates search results and sends them to the terminal. The input is the search result data, and the output is the data to be sent to the terminal. Specifically, the results are sent in JSON format.

[0177] Step 5:

[0178] The terminal displays the search results to the user. The input is the data received from the server, and the output is the display to the user. Specifically, the terminal displays the data using a GUI.

[0179] (Application example 1)

[0180] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0181] Modern manufacturing factory operations require automation and real-time progress management. In particular, there is a lack of technology that allows factory operators to give work instructions to robots via voice and efficiently manage their progress, resulting in a decline in production efficiency and work accuracy. There is also a need for a system that can quickly search for and provide necessary information.

[0182] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0183] In this invention, the server includes: means for a user to give instructions through voice input and for the terminal to convert the voice into text in real time; means for sending text data to the server and for the server to generate a summary and proposal of the information using a generative AI model; means for the terminal to display the summary and proposal sent from the server to the user; means for the terminal to convert the voice into text and then send the text to the server; means for the server to generate appropriate actions and information in response to the instructions and send the results to the terminal; means for the user to input progress and tasks by voice or text and for the terminal to send them to the server; means for the server to update the progress management database and generate the next action or alert; and means for the terminal to notify the user of the generated action or alert. This enables factory operators to give instructions to robots by voice and quickly manage progress and obtain related information based on those instructions.

[0184] A "user" is a person who uses the system to give instructions and input progress.

[0185] "Voice input" is a method in which a user gives instructions or information to a terminal by voice.

[0186] A "terminal" is a device that converts voice input from a user into text and transmits the text data to a server.

[0187] "Real-time" is a time characteristic that means processing data and returning results almost instantaneously.

[0188] "Means for converting to text" refers to technology or devices for converting voice data into character data.

[0189] "Text data" is a data format in which voice is converted into text.

[0190] A "server" is a central computer system that receives text data and performs analysis and information generation.

[0191] A "generative AI model" is an algorithm that uses artificial intelligence technology to analyze received data and generate information summaries and suggestions.

[0192] "Means of summarizing and suggesting information" is the process of using a generative AI model to extract important information from the data received and present the user with the next action to take.

[0193] The "means for displaying the summary and suggestions" refers to a technique by which the terminal presents the summary and suggestions sent from the server to the user visually or audibly.

[0194] "Progress" is data that indicates the progress of a particular task or project.

[0195] The "progress management database" is a database system for storing and managing progress information entered by the user.

[0196] "Means for generating actions and alerts" refers to the process for generating next actions and alerts based on the progress management database.

[0197] "Means for notifying actions and alerts" refers to techniques for notifying users of generated actions and alerts.

[0198] "Means for searching information" refers to the technology that allows a user to send information searched for by voice or text to a server and return the results.

[0199] The "means for displaying search results" refers to a technique for visually or audibly presenting the search results sent from the server to the user.

[0200] The system required to implement this invention includes functions such as voice input, real-time text conversion, generative AI model, progress management, information search, etc. These will be explained in detail below.

[0201] Hardware used

[0202] The present invention uses the following hardware:

[0203] Microphone: Collects voice input from the user.

[0204] Device: A device that converts speech to text and communicates with the server. This can be a smartphone, tablet, or computer.

[0205] Server: Runs the generative AI model, analyzes data, and manages progress.

[0206] Software used

[0207] The following software is used in the present invention:

[0208] speech_recognition: A Python library for converting speech to text.

[0209] requests: A Python library for sending HTTP requests.

[0210] Generative AI models: Artificial intelligence techniques for analyzing text and generating summaries and suggestions.

[0211] TextToSpeech: A library for converting text to speech.

[0212] System Operation

[0213] 1. The user speaks a command into the device's microphone, for example, "Start the next shaft inspection."

[0214] 2. The device converts the voice input into text in real time and sends the text data to the server.

[0215] 3. The server uses a generative AI model to analyze the received text data and generate appropriate actions and information, such as a summary and suggestion such as "Initiate shaft inspection protocol."

[0216] 4. The generated actions and information are sent to the terminal and notified to the user by display or voice.

[0217] 5. When the user inputs their progress via voice or text, the device sends the data to the server, which updates the progress management database. The next action or alert is generated and notified to the user via the device.

[0218] Specific examples

[0219] Consider a factory operator who wants to perform a shaft inspection on a production line. The operator gives a voice command such as "Start the next shaft inspection." The device converts the voice to text and sends it to the server. The server uses a generative AI model to analyze the message and return it to the device: "Starting shaft inspection protocol." The operator proceeds with the work according to the protocol and reports the progress by voice. The server updates the progress management database and notifies the operator of the next task or alert.

[0220] Prompt Sentence Examples

[0221] Here are some examples of specific prompts:

[0222] "We will begin testing Shaft A next week."

[0223] "What are the technical specifications of the new product?"

[0224] This system allows factory operators to give voice instructions, which enable progress management and the rapid acquisition of related information.

[0225] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0226] Step 1:

[0227] The user inputs voice instructions through the microphone of the terminal, and the input voice data is sent to the terminal.

[0228] Step 2:

[0229] The device receives the voice data and converts it into text data in real time. Specifically, it uses the speech_recognition library to analyze the voice and output it as text data.

[0230] Step 3:

[0231] The text data generated by the terminal is sent to the server using an HTTP request.

[0232] Step 4:

[0233] The server inputs the received text data into a generative AI model for data analysis. The AI ​​model analyzes the instructions based on the prompt and generates a summary and proposal. For example, in response to the instruction "Start the next shaft inspection," it generates the summary "Start the shaft inspection protocol."

[0234] Step 5:

[0235] The server sends the generated summary and proposal to the terminal as an HTTP response.

[0236] Step 6:

[0237] The device receives the summary and suggestions from the server and displays them to the user, either by displaying the summary as text on the screen or by announcing it aloud using the TextToSpeech library.

[0238] Step 7:

[0239] The user inputs progress by voice or text, and the data is received by the terminal. For example, the user may report by voice that "the shaft inspection is complete."

[0240] Step 8:

[0241] The device converts the voice data into text in real time and sends the text data to the server, again using an HTTP request.

[0242] Step 9:

[0243] The server updates the progress management database based on the received text data. For example, it moves "Shaft Inspection" from the in-progress task list to the completed task list.

[0244] Step 10:

[0245] Based on the results of the updates to the progress management database, the server generates the next action or alert, based on pre-defined rules and conditions.

[0246] Step 11:

[0247] The server generates actions and alerts and sends them to the terminal as HTTP responses.

[0248] Step 12:

[0249] The device notifies the user of actions and alerts from the server, either by displaying them on the device screen or by voice using the TextToSpeech library.

[0250] Step 13:

[0251] When a user wants to search for information, the user inputs a query by voice or text, and the device receives the input. For example, the user may input a query such as "What are the technical specifications of a new product?"

[0252] Step 14:

[0253] The device converts the voice data into text and sends the text data to the server, again using an HTTP request.

[0254] Step 15:

[0255] The server searches the database based on the text data to retrieve relevant information, then analyzes the retrieved information using a generative AI model to generate search results in a format appropriate for the user.

[0256] Step 16:

[0257] The search results generated by the server are sent to the terminal as an HTTP response.

[0258] Step 17:

[0259] The device displays the search results from the server to the user, either as text on the screen or as audio using the TextToSpeech library.

[0260] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0261] Understood. Below is the "Form for carrying out the invention".

[0262] The sales support AI app of this invention is a system that converts user voice input into text in real time and uses that data to summarize information and make suggestions using generative AI. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and by adjusting the content of suggestions based on the user's emotional state, it provides more effective sales support.

[0263] Business negotiation support

[0264] During business negotiations, the user launches the app and uses voice input. The device converts the voice input into text in real time using voice recognition technology and sends the text data to the server. The server then uses generative AI to analyze the text data, summarizes key points, and suggests next steps. The server then uses an emotion engine to analyze the user's emotions from the voice data and adjusts the proposal content based on those emotions. The adjusted summary and proposal are then displayed to the user via the device.

[0265] For example, if a user is introducing a new product during a sales meeting and the customer asks, "How much does this product cost?", the device will recognize the voice, convert it into text data, and send it to the server. The server will analyze the text data and recognize the user's emotions through an emotion engine. If the user seems nervous, it will generate a more polite explanation and return it to the device. As a result, the user can respond, "The price of this product is XX yen, but we are running a discount campaign for a limited time."

[0266] Progress management support

[0267] To update progress, the user inputs progress information by voice or text. The device performs voice recognition or text analysis on this information and sends it to the server as text data. The server updates the progress management database and generates the next action or alert. The server also uses an emotion engine to analyze the user's emotions when entering progress information, and generates a faster alert if the information is urgent or important. The generated action or alert is then notified to the user via the device.

[0268] For example, if a user types, "I need to set up my next meeting with Client A," the device recognizes this, converts it into text data, and sends it to the server. The server updates the progress management database, analyzes the user's state of tension using an emotion engine, and then generates a high-priority alert and sends it to the device. As a result, the user receives a real-time notification that "A meeting with Client A has been set up."

[0269] Information search support

[0270] In information searches, users input the information they are looking for by voice or text. The device performs voice recognition or text analysis on the input and sends the text data to a server. The server receives the query and searches a database for relevant information. The emotion engine adjusts the content of the search results to take into account the user's emotions. For example, if the user is in a hurry, the device will display the search results in a concise summary format. The device then displays the results to the user and provides voice feedback as needed.

[0271] For example, if a user says, "Tell me the technical specifications of this product," the device converts this into text data and sends it to the server. The server searches the database for the relevant technical specifications, and if the emotion engine detects the user's impatience, it generates a concise and easy-to-understand technical specification and returns it to the device. As a result, the user can quickly and accurately tell the user, "The technical specifications of this product are ____, and for more details, please see this link."

[0272] In this way, by incorporating an emotion engine, it is possible to respond according to the user's emotional state, further improving the quality and efficiency of sales activities.

[0273] The processing flow will be explained below.

[0274] Understood. Below, I will explain in detail the processing steps of the sales support AI app that combines an emotion engine.

[0275] Business negotiation support

[0276] Step 1:

[0277] The user launches the app by tapping the app icon and begins voice input.

[0278] Step 2:

[0279] The device activates the voice input function and records the user's voice in real time, which is then sent to a voice recognition algorithm.

[0280] Step 3:

[0281] The device recognizes the recorded voice and converts it into text data, which is then temporarily saved.

[0282] Step 4:

[0283] The device sends text and voice data to the server in real time, where it is uploaded via a network connection.

[0284] Step 5:

[0285] The server receives the text data and uses generative AI to summarize the main points, while simultaneously analyzing the audio data with an emotion engine to recognize the user's emotional state.

[0286] Step 6:

[0287] Based on the analysis, the server generates suggestions that match the user's emotions, for example, if the user is nervous, a suggestion with more polite explanations will be generated.

[0288] Step 7:

[0289] The server then sends the generated summary data and sentiment-based suggestions to the device, where the data is transferred over the network.

[0290] Step 8:

[0291] The device displays summary data and suggestions to the user, providing on-screen text and, optionally, audio feedback.

[0292] Progress management support

[0293] Step 1:

[0294] The user enters progress information into the app, using voice or text input functionality to enter progress details.

[0295] Step 2:

[0296] The device converts voice into text in real time and temporarily stores the input text data.

[0297] Step 3:

[0298] The device sends the text data to the server, where it is uploaded over a network connection.

[0299] Step 4:

[0300] The server receives the progress data and updates the progress management database, and the new progress information is recorded in the database.

[0301] Step 5:

[0302] The server generates the next action or alert based on the progress data, using an emotion engine to analyze the user's emotional state and evaluate the urgency and importance of the action or alert.

[0303] Step 6:

[0304] The server sends the generated actions and alerts to the terminal, and the data is transferred over the network.

[0305] Step 7:

[0306] The device notifies the user of the generated actions and alerts by displaying them on the screen and / or providing audio feedback.

[0307] Information search support

[0308] Step 1:

[0309] The user enters the information they want to search for by voice or text, taps the search button, and enters search keywords.

[0310] Step 2:

[0311] The device converts voice input into text in real time and temporarily stores the text data.

[0312] Step 3:

[0313] The device sends text data to the server, where it is uploaded via the network.

[0314] Step 4:

[0315] The server searches the database based on the received query and extracts data that matches the specified keywords.

[0316] Step 5:

[0317] The server generates search results and analyzes the user's emotional state using an emotion engine, and adjusts the display format of the search results based on the user's emotion.

[0318] Step 6:

[0319] The server sends the adjusted search results to the device, and the data is transferred over the network.

[0320] Step 7:

[0321] The device displays the search results to the user, providing text on the screen and, optionally, audio feedback.

[0322] By following these steps, a sales support AI app incorporating an emotion engine will provide responses that correspond to the user's emotional state, improving the quality and efficiency of sales activities.

[0323] Example 2

[0324] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0325] In the processing of voice input in sales activities, the challenge is to generate appropriate information summaries and suggestions that take the user's emotions into account, thereby improving the efficiency of progress management and information retrieval. Conventional systems can only respond uniformly, ignoring the user's emotions, and it is difficult to respond flexibly according to the user's situation.

[0326] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0327] In this invention, the server includes means for analyzing the user's emotions from voice data and adjusting the content of proposals based on the emotions, means for analyzing the user's emotions when entering progress and generating a prompt alert if the urgency or importance is high, and means for adjusting the display content in consideration of the user's emotional state when generating search results. This enables flexible responses according to the user's emotional state, improving the quality and efficiency of sales activities.

[0328] Understood. Below are definitions of important words.

[0329] "Voice input" refers to the process by which a user inputs voice through a microphone.

[0330] "Convert to text" refers to the process of converting audio data into written information in real time.

[0331] "Generative AI" refers to artificial intelligence technology that learns from large amounts of data and generates new information and suggestions.

[0332] "Emotion analysis" refers to the process of analyzing a user's emotional state from input voice and text data.

[0333] "Emotion engine" refers to software or hardware used to recognize the emotional state of a user.

[0334] "Adjusting suggestions" refers to the process of changing the content of generated suggestions depending on the user's emotional state.

[0335] "Progress information" refers to data that indicates the progress or task status related to a project or sales activity.

[0336] "Progress management database" refers to a database for managing and storing progress information.

[0337] "Generating an alert" refers to the process of creating a notification to draw attention to an event of high urgency or importance.

[0338] "Search query" means a command or question entered by a user into a system in search of specific information.

[0339] "Search Results" refers to relevant information retrieved from a database based on a user's search query.

[0340] "Adjusting display content" refers to the process of changing the format and detail of the information displayed to take into account the user's emotional state.

[0341] MODE FOR CARRYING OUT THE INVENTION

[0342] The sales support system of the present invention converts a user's voice input into text in real time, and uses generative AI to summarize the information and make suggestions based on that data. It also incorporates an emotion engine that recognizes the user's emotions, and can adjust the content of suggestions based on the user's emotional state. A specific embodiment of this system is described below.

[0343] Hardware and software used

[0344] This system uses the following hardware and software:

[0345] Devices (e.g. smartphones, tablets, laptops)

[0346] Server (cloud server or on-premise server)

[0347] Voice recognition technology (e.g., Google Cloud Speech-to-Text)

[0348] Generative AI models (e.g., OpenAI's GPT-3)

[0349] Sentiment analysis engine (e.g., Microsoft Azure's Text Analytics API)

[0350] Specific Example of the System

[0351] Business negotiation support

[0352] 1. The user launches the app and uses voice input, for example, "What is the price of the new product?"

[0353] 2. The device converts the voice input into text in real time using Google Cloud Speech-to-Text technology and sends the text data to the server.

[0354] 3. The server analyzes the received text data using a generative AI model (GPT-3), extracts key points, and suggests the next action. For example, it generates a suggestion such as, "The price of this product is XX yen."

[0355] 4. The server uses an emotion analysis engine to analyze the user's emotions from the voice data, and if the user is nervous, adjusts the explanation or suggestions to be more friendly.

[0356] 5. The adjusted summary and suggestions are displayed to the user via the device.

[0357] Example prompt sentence:

[0358] "New product pricing information that can be used in business negotiations"

[0359] Progress management support

[0360] 1. The user inputs status information by voice or text, for example, "I need to schedule my next meeting with Client A."

[0361] 2. The device uses voice recognition technology to convert the voice into text and sends the data to the server.

[0362] 3. The server updates the progress management database and generates the next action or alert, for example, "A meeting has been scheduled with Client A."

[0363] 4. The server uses an emotion analysis engine to analyze the user's emotions when entering progress, and if it determines that the situation is urgent, it generates a prompt alert.

[0364] 5. The device notifies the user of the generated action or alert.

[0365] Example prompt sentence:

[0366] "Schedule my next meeting"

[0367] Information search support

[0368] 1. The user enters the information they are looking for by voice or text, for example, "What are the technical specifications for this product?"

[0369] 2. The device uses voice recognition technology to convert the voice into text and sends the data to the server.

[0370] 3. The server searches the database for relevant information and generates search results, such as "The technical specifications of this product are ____."

[0371] 4. When generating search results, the server analyzes the user's emotional state using an emotion analysis engine and adjusts the displayed content accordingly. For example, if the user is anxious, the server will provide information in a concise format.

[0372] 5. The device displays the adjusted search results to the user and provides audio feedback if necessary.

[0373] Example prompt sentence:

[0374] "What are the technical specifications of the product?"

[0375] This system aims to improve the quality and efficiency of sales activities by enabling flexible responses according to the user's emotional state.

[0376] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0377] Sales Support Processing Steps

[0378] Step 1:

[0379] The user launches the app and performs voice input.

[0380] Specifically, the user opens the app on their smartphone and asks, "What is the price of the new product?"

[0381] Input: User's voice data

[0382] Output: Audio input data

[0383] Step 2:

[0384] The device converts voice input into text in real time using speech recognition technology (e.g., Google Cloud Speech-to-Text).

[0385] Specifically, the device's microphone captures audio and converts it into text via a cloud service.

[0386] Input: Voice input data

[0387] Output: Text data

[0388] Step 3:

[0389] The terminal transmits the converted text data to the server.

[0390] Specifically, the device sends text data to the server over a secure HTTPS connection.

[0391] Input: Text data

[0392] Output: The text data sent

[0393] Step 4:

[0394] The server uses a generative AI model (e.g., GPT-3) to analyze the text data, extract key points, and suggest the next action.

[0395] Specifically, the server passes the text data to the analysis engine, and the generative AI model generates information about the "price of the new product."

[0396] Input: Text data sent

[0397] Output: Summary of proposal

[0398] Step 5:

[0399] The server uses a sentiment analysis engine (for example, Microsoft Azure's Text Analytics API) to analyze the user's emotions from the voice data and adjusts the suggestions based on those emotions.

[0400] Specifically, the server transfers voice data to an analysis engine, and if it detects that the user is in a tense state, it changes the suggestions to be more polite.

[0401] Input: Audio data, summarized proposal

[0402] Output: Sentiment-adjusted recommendations

[0403] Step 6:

[0404] The server sends the adjusted summary and suggestions to the terminal.

[0405] Specifically, the adjusted proposal is sent to the device over a secure connection.

[0406] Input: Emotion-adjusted recommendations

[0407] Output: Adjustment proposals submitted

[0408] Step 7:

[0409] The device displays the adjusted proposal to the user.

[0410] Specifically, the device will display on the screen, "The price of this product is XX yen. We are running a discount campaign for a limited time."

[0411] Input: Adjustment proposal submitted

[0412] Output: The suggestions displayed to the user

[0413] Progress management support processing steps

[0414] Step 1:

[0415] The user enters progress information by voice or text.

[0416] Specifically, the user might say or type in text, "I need to schedule my next meeting with Client A."

[0417] Input: Audio or text data

[0418] Output: Progress input data

[0419] Step 2:

[0420] The device converts the voice input into text using voice recognition technology and sends the data to the server.

[0421] Specifically, the terminal converts the progress input into text and securely transmits it to the server.

[0422] Input: Progress input data

[0423] Output: Progress data sent

[0424] Step 3:

[0425] The server updates the progress management database and generates the next action or alert.

[0426] As a specific operation, the server adds new progress information to the database and generates an action as "meeting set up."

[0427] Input: Submitted progress data

[0428] Output: Updated progress data, generated actions

[0429] Step 4:

[0430] The server uses an emotion analysis engine to analyze the user's emotions when entering progress information, and generates a prompt alert if the information is of high urgency or importance.

[0431] Specifically, the server analyzes the emotions expressed when entering progress, and if it determines that the situation is urgent, it generates a high-priority alert.

[0432] Input: Progress data, user emotion data

[0433] Output: The generated alert

[0434] Step 5:

[0435] The terminal notifies the user of the generated actions and alerts.

[0436] Specifically, the device will display a notification on the screen saying, "A meeting with Client A has been set up."

[0437] Input: Generated action, alert

[0438] Output: Information reported to the user

[0439] Information retrieval support processing steps

[0440] Step 1:

[0441] The user inputs the information they are looking for by voice or text.

[0442] Specifically, the user says, "Tell me the technical specifications of this product."

[0443] Input: Audio or text data

[0444] Output: Search query

[0445] Step 2:

[0446] The device converts the input into text using voice recognition technology and sends the data to the server.

[0447] Specifically, the device converts the voice input into text and sends it to the server.

[0448] Input: search query

[0449] Output: Submitted search query data

[0450] Step 3:

[0451] The server receives the query and retrieves the relevant information from a database.

[0452] Specifically, the server queries a product database to obtain the relevant technical specification information.

[0453] Input: Submitted search query data

[0454] Output: Related information

[0455] Step 4:

[0456] When generating search results, the server analyzes the user's emotional state using an emotion analysis engine and adjusts the display content.

[0457] Specifically, if the server determines that the user is in a hurry, it displays the technical specifications in a concise and easy-to-understand format.

[0458] Input: Related information, user emotion data

[0459] Output: Search results tailored based on sentiment

[0460] Step 5:

[0461] The device displays the adjusted search results to the user and provides audio feedback if necessary.

[0462] Specifically, the device will display on the screen, "The technical specifications of this product are XX. Please refer to this link for details," and the voice assistant will also read it out loud.

[0463] Input: Refined search results

[0464] Output: Search results displayed to the user

[0465] (Application example 2)

[0466] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0467] Sales activities in brick-and-mortar stores require immediate responses to customer questions and reactions, but conventional systems have difficulty taking customer emotions into account, often failing to provide appropriate information or suggestions. Furthermore, in situations where salespeople are at a loss as to how to respond, a system is needed that can generate and provide appropriate suggestions in real time. The present invention aims to solve these problems and provide a system that supports salespeople in effectively responding to customers.

[0468] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for using an emotion analysis engine that analyzes emotions from audio data and video data, means for generating summaries and proposals using generative AI, and means for the server to transmit summaries and proposals adjusted based on the emotion analysis. This makes it possible to provide information and proposals in real time that take customer emotions into consideration.

[0469] "User" refers to an individual or entity who uses the System to provide voice or text input.

[0470] A "terminal" is a device operated by a user that has the function of sending voice input and text data to a server.

[0471] "Server" refers to a central processing unit that receives text data sent from a device, processes the information using generative AI or an emotion analysis engine, and returns the results to the device.

[0472] "Generative AI" refers to an artificial intelligence engine that analyzes text data and automatically generates summaries of information and suggestions.

[0473] An "emotion analysis engine" is an engine that has the ability to analyze a user's emotions from audio and video data and adjust the content of suggestions based on that information.

[0474] "Voice input" refers to user voice data collected by the terminal, which is converted into text in real time.

[0475] "Text data" refers to information in which voice input is converted into text in real time.

[0476] A "summary" refers to information that is concisely summarized by generative AI that analyzes text data and extracts important points.

[0477] "Suggestion" refers to the next action or information the user should take, generated by generative AI based on text data.

[0478] "Progress management database" refers to a database for storing and managing user progress information.

[0479] An "alert" refers to urgent or important information that is notified to the user when the progress management database is updated.

[0480] "Search results" refers to information that the server searches for related information from the database and presents to the user.

[0481] The present invention is a system that supports sales activities in brick-and-mortar stores, and makes it possible to provide information and suggestions in real time based on the user's voice input and video data.

[0482] System Overview

[0483] The system includes the following main elements:

[0484] 1. User (salesperson): Operates the system and provides voice input.

[0485] 2. Terminal: A device operated by the user (e.g., smart glasses) that collects voice input and sends it to a server.

[0486] 3. Server: Analyzes data sent from the device and processes the information using generative AI and an emotion analysis engine.

[0487] Hardware and Software Details

[0488] Smart glasses: Devices with AR capabilities and microphones (e.g., Google Glass) that collect audio and visual data from the user.

[0489] Speech Recognition API: Uses the Google Cloud Speech-to-Text API to convert voice input to text in real time.

[0490] Generative AI model: Uses OpenAI GPT to analyze text data and generate information summaries and suggestions.

[0491] Sentiment analysis engine: Uses the Microsoft Azure Emotion API to analyze emotions from audio and video data.

[0492] Cloud server: Data processing is performed using Amazon Web Services (AWS).

[0493] Operation explanation

[0494] Receiving audio input

[0495] A user (salesperson) puts on the smart glasses during a sales meeting and starts a conversation with a customer. The microphone in the smart glasses collects the salesperson's voice input and temporarily stores the voice data. The stored voice data is converted into text data in real time by the Google Cloud Speech-to-Text API.

[0496] Sending and analyzing text data

[0497] The converted text data is sent to a cloud server (AWS). A generative AI model (OpenAI GPT) on the cloud server analyzes the text data, summarizes key points, and generates next actions and suggestions. In parallel, the audio and video data is analyzed by the Microsoft Azure Emotion API to determine the customer's emotions. Based on the results of the emotion analysis, the suggestions generated by the generative AI are adjusted appropriately.

[0498] Viewing Proposals

[0499] The final tailored summary and recommendations are then displayed on the smart glasses display, allowing the salesperson to respond appropriately to the customer.

[0500] Specific examples

[0501] For example, if a customer asks a salesperson, "How much does this product cost?", the smart glasses will collect the voice and convert it into text data in real time. This text data is sent to a server and analyzed by a generative AI model. The sentiment analysis engine will analyze the customer's emotions, and if the customer seems impatient, a short, concise, and unassuming answer will be generated. As a result, the salesperson can respond, "The price of this product is XX yen, but we are running a discount campaign for a limited time."

[0502] Prompt Sentence Examples

[0503] Convert user questions into text in real time, analyze customer sentiment, and generate optimal responses.

[0504] Example input:

[0505] Salesperson: "Look at this product, it has special features."

[0506] Customer: "What's the price?"

[0507] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0508] Step 1:

[0509] A user (salesperson) puts on smart glasses and starts a sales negotiation. As input, the microphone in the smart glasses collects voice data of the conversation with the customer. As output, the voice data is temporarily stored in the built-in memory. Specifically, the smart glasses continuously capture voice through the microphone.

[0510] Step 2:

[0511] The collected voice data is converted into text data in real time. As input, the voice data collected in step 1 is sent to the Google Cloud Speech-to-Text API. As output, the voice data is converted into text data. Specifically, the device (smart glasses) calls the Google Cloud Speech-to-Text API to convert the voice data into text.

[0512] Step 3:

[0513] The converted text data is sent to the server. As input, the text data converted in step 2 is sent to the cloud server (AWS). As output, the text data is saved on the server. In concrete terms, the device sends the text data to the cloud server via the Internet.

[0514] Step 4:

[0515] The server analyzes the text data and generates a summary and suggestions based on key points. The text data sent to the server in step 3 is received as input. The summary and suggestions are generated as output. Specifically, a generative AI model (OpenAI GPT) on the server analyzes the text data, extracts key elements, and generates suggestions.

[0516] Step 5:

[0517] In parallel, the audio and video data are sent to the emotion analysis engine on the server for analysis. As input, the audio data collected in step 1 and the video data captured by the smart glasses camera are sent to the server. As output, customer emotion data is generated. Specifically, the server uses the Microsoft Azure Emotion API to analyze emotions from the audio and video.

[0518] Step 6:

[0519] The generated proposals are adjusted based on the results of the sentiment analysis engine. The summaries and proposals generated in step 4 and the emotion data generated in step 5 are used as input. The output is a summary and proposal adjusted based on emotion. Specifically, the generative AI model on the server uses the emotion data to optimize the summaries and proposals.

[0520] Step 7:

[0521] The final summary and suggestions are sent to the terminal and displayed to the user. As input, the adjusted summary and suggestions generated in step 6 are sent from the cloud server to the terminal. As output, the suggestions are displayed on the display of the smart glasses. Specifically, the server sends the summary and suggestions to the smart glasses via the Internet, and the smart glasses display them.

[0522] Prompt Sentence Examples

[0523] Convert user questions into text in real time, analyze customer sentiment, and generate optimal responses.

[0524] Example input:

[0525] Salesperson: "Look at this product, it has special features."

[0526] Customer: "What's the price?"

[0527] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0528] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0529] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0530] [Second embodiment]

[0531] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0532] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0533] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0534] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0535] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0536] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0537] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0538] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0539] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0540] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0541] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0542] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0543] Understood. Below is the "Form for carrying out the invention".

[0544] In one embodiment of the invention, the system begins with a user providing instructions through voice input, which the device converts in real time into text and sends to the server. The server then uses generative AI to analyze the text data, generating summaries and suggestions, which are then sent back to the device for display to the user. When the user inputs progress, this is recorded in a progress management database, and the user is notified of next actions or alerts. Furthermore, when the user searches for specific information, the server searches for related data and provides the results to the user.

[0545] Business negotiation support

[0546] When a user launches the app, the app activates its voice input function. The device converts the user's speech into text in real time and sends the text data to a server. The server then uses generative AI to analyze the text data, summarize key points, and suggest next steps. These summaries and suggestions are displayed to the user on the device.

[0547] For example, suppose a user is introducing a new product during a business meeting and the client asks, "What is the price of this product?" In this case, the device recognizes the voice as a "price question" and sends the text data to the server. The server analyzes past data, generates appropriate price information, and returns it to the device. As a result, the user can respond to the client by saying, "The price of this product is XX yen."

[0548] Progress management support

[0549] In progress management, the user inputs the progress of each business negotiation or task by voice or text. The device recognizes this and sends the data to the server. The server updates the progress management database and generates the next action or alert. These actions and alerts are notified to the user via the device.

[0550] For example, when a user says, "Set up a meeting with Client A next week," the device converts the voice to text and sends it to the server. The server updates the progress management database and generates an alert to set up the next meeting. The device notifies the user when the meeting date approaches.

[0551] Information search support

[0552] With the information search function, users search for the information they need using voice or text. The device analyzes this input and sends the query to the server. The server then searches for relevant information in a database, generates results, and sends them to the device. The user can then view the search results through the device.

[0553] For example, if a user says, "I want to know the technical specifications of a new product," the device converts the instruction into text and sends it to the server. The server then searches the database for the new product's technical specifications and returns them to the device. The user can then view the specifications on the screen.

[0554] As a result, the system standardizes the quality of business negotiations, improves the efficiency of progress management, and enables quick and accurate information acquisition.

[0555] The processing flow will be explained below.

[0556] Understood. Below are the specific processing steps.

[0557] Business negotiation support

[0558] Step 1:

[0559] The user launches the app by tapping the app icon.

[0560] Step 2:

[0561] The terminal activates the voice input function and starts recording the user's voice in real time.

[0562] Step 3:

[0563] The device converts the recorded audio into text data in real time, using a speech recognition algorithm to output the audio as text.

[0564] Step 4:

[0565] The terminal transmits the text data to the server, and the text data is uploaded to the server via the network.

[0566] Step 5:

[0567] The server receives the text data and uses generative AI to summarize the information, including important points and next actions.

[0568] Step 6:

[0569] The server generates information summaries and suggestions, and creates optimal suggestions based on the results of analyzing the text data.

[0570] Step 7:

[0571] The server transmits the generated summary data and the proposal to the terminal, and transmits the data to the terminal via the network.

[0572] Step 8:

[0573] The device displays summary data and suggestions to the user, either as text on the screen or as audio feedback.

[0574] Progress management support

[0575] Step 1:

[0576] The user types a progress update into the app or gives a voice command. They tap the progress update button and enter the information.

[0577] Step 2:

[0578] The terminal performs voice recognition or text analysis on the input content to generate text data.

[0579] Step 3:

[0580] The device sends the text data to the server, where it is uploaded via the network.

[0581] Step 4:

[0582] The server receives the progress data and updates the progress management database, recording the new progress information.

[0583] Step 5:

[0584] The server analyzes the progress and generates the next action or alert: Create the necessary tasks or alerts.

[0585] Step 6:

[0586] The server sends the generated tasks and alerts to the terminal and returns the data via the network.

[0587] Step 7:

[0588] The device notifies the user of tasks and alerts, displaying them on the screen and providing audio feedback where appropriate.

[0589] Information search support

[0590] Step 1:

[0591] When a user wants to search for specific information, they input it by voice or text, then tap the search button.

[0592] Step 2:

[0593] The terminal performs voice recognition or text analysis on the input content to generate text data.

[0594] Step 3:

[0595] The device sends the text data to the server, where it is uploaded via the network.

[0596] Step 4:

[0597] The server receives the query and searches the database to extract information that matches the specified keywords.

[0598] Step 5:

[0599] The server generates search results and sends them to the device, organizes the results, and sends the data to the device.

[0600] Step 6:

[0601] The device displays the search results to the user, displaying them on the screen and providing audio feedback if necessary.

[0602] By following the above steps, this system provides functions such as support for business negotiations, progress management, and information search.

[0603] Example 1

[0604] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0605] Conventional sales negotiation support systems, progress management systems, and information search systems each function independently, resulting in low user convenience. Furthermore, there is a lack of systems that can efficiently convert voice input into text and then centrally manage the subsequent analysis, proposals, progress management, and information search. This creates challenges for users, making it difficult to standardize the quality of sales negotiations, efficiently manage progress, and quickly obtain information.

[0606] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0607] In this invention, the server includes means for analyzing information using a generative artificial intelligence model and generating summaries and proposals, means for updating the progress management database and generating next actions and alerts, and means for searching the database for related information, generating search results, and sending them to the terminal. This allows for centralized management of business negotiation support, progress management, and information search, enabling users to efficiently give instructions and obtain information.

[0608] "User" refers to any individual or corporation that uses this system.

[0609] "Voice input" refers to a means by which a user gives instructions to a system using voice.

[0610] "Terminal" refers to a hardware device that converts speech to text, transmits the data to a server, and displays the results.

[0611] "Server" refers to a computer system that uses generative artificial intelligence models to analyze text data and generate summaries, suggestions, progress management, and information search results.

[0612] A "generative artificial intelligence model" refers to an algorithm or program that analyzes input text data and generates summaries or suggestions based on it.

[0613] "Text data" refers to data resulting from converting voice input into text.

[0614] "Analyzing information" refers to the process of understanding input text data and generating summaries or suggestions based on that content.

[0615] A "summary" refers to a concise summary of the important points of text data.

[0616] "Suggestion" refers to showing the user the next action or countermeasure to be taken based on the analysis results.

[0617] A "progress management database" refers to a database for recording and managing the progress of business negotiations and tasks.

[0618] "Action" refers to the specific next steps or actions to be taken.

[0619] "Alert" refers to a warning or reminder that notifies the user of important information or progress.

[0620] "Information search" refers to a search performed by a user to obtain specific information.

[0621] "Database" refers to a data storage system in which related information is stored.

[0622] "Search Results" refers to a server-generated answer or collection of data in response to an information search query.

[0623] This invention is a system that unifies the management of business negotiation support, progress management, and information search. This system is composed of three main components: users, terminals, and servers.

[0624] Business negotiation support

[0625] When a user launches the app, the device activates the voice input function. When the user speaks, the device converts the speech into text in real time and sends the text data to the server. The server then analyzes the text data using a generative artificial intelligence model (e.g., OpenAI's GPT series), summarizes the key points, and suggests the next action to take. The suggested summary and action are displayed to the user on the device.

[0626] For example, if a user is introducing a new product and the client asks, "What is the price of this product?", the device converts the speech into text "Question about price" and sends it to the server. The server analyzes past data, generates appropriate price information, and returns it to the device. The user can then reply to the client, "The price of this product is XX yen."

[0627] Progress management

[0628] When the user inputs progress by voice or text, the device sends the data to the server, which updates the progress management database (e.g., PostgreSQL) and generates the next action or alert. The generated action or alert is then notified to the user via the device.

[0629] For example, if a user says, "Set up a meeting with Client A next week," the device converts the speech into text and sends it to the server. The server updates the progress management database and generates an alert to set up the next meeting. When the meeting date approaches, the device notifies the user.

[0630] Information Search

[0631] When a user searches for the information they need by voice or text, the device sends the query to the server, which then searches for relevant information in a database (e.g., Elasticsearch), generates results, and sends them to the device. The user can then view the search results through the device.

[0632] For example, if a user says, "I want to know the technical specifications of a new product," the device converts the instruction into text and sends it to the server. The server then searches the database for the new product's technical specifications and returns them to the device. The user can then view the specifications on the device's screen.

[0633] As a result, this system can standardize the quality of business negotiations, improve the efficiency of progress management, and enable quick and accurate information acquisition.

[0634] Prompt Sentence Examples

[0635] For sales support: "Please summarize the notes from your meeting with Client X and tell me what action to take next."

[0636] For progress management support: "How is project Y progressing? What's the next action?"

[0637] For information search assistance: "Please tell me the technical specifications for new product Z."

[0638] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0639] Understood. Below we will explain in detail the processing flow of the system program, divided into processing steps.

[0640] Business negotiation support

[0641] Step 1:

[0642] A user launches an app and activates the voice input function. The input is the command to launch the app, and the output is that voice input is enabled.

[0643] Step 2:

[0644] The user provides voice input, which the device converts into text in real time. The input is the user's voice data, and the output is text data. Specifically, the device uses voice recognition software (e.g., Google Speech-to-Text API).

[0645] Step 3:

[0646] The terminal sends the generated text data to the server. The input is the text data, and the output is the data sent to the server. Specifically, the terminal sends data to the server using an HTTPS request.

[0647] Step 4:

[0648] The server uses a generative AI model to analyze the transmitted text data. The input is text data, and the output is a summary and proposed data. Specifically, the server uses a generative AI model such as GPT-4 to perform the analysis and generation process.

[0649] Step 5:

[0650] The server sends the generated summary and proposal to the terminal. The input is the summary and proposal data, and the output is the data to be sent to the terminal. Specifically, the server sends the generated results in JSON format.

[0651] Step 6:

[0652] The terminal displays the received summary and suggestions to the user. The input is the data received from the server, and the output is the visualization to the user. Specifically, the terminal displays the data using a GUI.

[0653] Progress management

[0654] Step 1:

[0655] The user inputs their progress by voice or text. The input is the user's voice or text data, and the output is text data. Specifically, the device recognizes the voice and converts it into text.

[0656] Step 2:

[0657] The terminal sends the text data to the server. The input is the text data, and the output is the data sent to the server. Specific operations use an HTTPS request.

[0658] Step 3:

[0659] The server updates the progress management database. The input is text data, and the output is the updated database. Specifically, the server updates the database using an SQL query.

[0660] Step 4:

[0661] The server generates the next action or alert. The input is the updated database information, and the output is the action / alert data. Specifically, the server checks for pending tasks and generates the next action or alert.

[0662] Step 5:

[0663] The device notifies the user of the generated actions and alerts. The input is the action / alert data sent from the server, and the output is the notification to the user. Specifically, the device performs push notifications and in-app notifications.

[0664] Information Search

[0665] Step 1:

[0666] The user searches for the information they need by voice or text. The input is the user's query (voice or text), and the output is the query data. Specifically, the device performs voice recognition and converts it into text.

[0667] Step 2:

[0668] The device sends the query to the server. The input is the query data, and the output is the data sent to the server. Specific operations use an HTTPS request.

[0669] Step 3:

[0670] The server searches for relevant information from a database. The input is the query data, and the output is the search result data. Specifically, the server uses a search engine such as Elasticsearch.

[0671] Step 4:

[0672] The server generates search results and sends them to the terminal. The input is the search result data, and the output is the data to be sent to the terminal. Specifically, the results are sent in JSON format.

[0673] Step 5:

[0674] The terminal displays the search results to the user. The input is the data received from the server, and the output is the display to the user. Specifically, the terminal displays the data using a GUI.

[0675] (Application example 1)

[0676] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0677] Modern manufacturing factory operations require automation and real-time progress management. In particular, there is a lack of technology that allows factory operators to give work instructions to robots via voice and efficiently manage their progress, resulting in a decline in production efficiency and work accuracy. There is also a need for a system that can quickly search for and provide necessary information.

[0678] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0679] In this invention, the server includes: means for a user to give instructions through voice input and for the terminal to convert the voice into text in real time; means for sending text data to the server and for the server to generate a summary and proposal of the information using a generative AI model; means for the terminal to display the summary and proposal sent from the server to the user; means for the terminal to convert the voice into text and then send the text to the server; means for the server to generate appropriate actions and information in response to the instructions and send the results to the terminal; means for the user to input progress and tasks by voice or text and for the terminal to send them to the server; means for the server to update the progress management database and generate the next action or alert; and means for the terminal to notify the user of the generated action or alert. This enables factory operators to give instructions to robots by voice and quickly manage progress and obtain related information based on those instructions.

[0680] A "user" is a person who uses the system to give instructions and input progress.

[0681] "Voice input" is a method in which a user gives instructions or information to a terminal by voice.

[0682] A "terminal" is a device that converts voice input from a user into text and transmits the text data to a server.

[0683] "Real-time" is a time characteristic that means processing data and returning results almost instantaneously.

[0684] "Means for converting to text" refers to technology or devices for converting voice data into character data.

[0685] "Text data" is a data format in which voice is converted into text.

[0686] A "server" is a central computer system that receives text data and performs analysis and information generation.

[0687] A "generative AI model" is an algorithm that uses artificial intelligence technology to analyze received data and generate information summaries and suggestions.

[0688] "Means of summarizing and suggesting information" is the process of using a generative AI model to extract important information from the data received and present the user with the next action to take.

[0689] The "means for displaying the summary and suggestions" refers to a technique by which the terminal presents the summary and suggestions sent from the server to the user visually or audibly.

[0690] "Progress" is data that indicates the progress of a particular task or project.

[0691] The "progress management database" is a database system for storing and managing progress information entered by the user.

[0692] "Means for generating actions and alerts" refers to the process for generating next actions and alerts based on the progress management database.

[0693] "Means for notifying actions and alerts" refers to techniques for notifying users of generated actions and alerts.

[0694] "Means for searching information" refers to the technology that allows a user to send information searched for by voice or text to a server and return the results.

[0695] The "means for displaying search results" refers to a technique for visually or audibly presenting the search results sent from the server to the user.

[0696] The system required to implement this invention includes functions such as voice input, real-time text conversion, generative AI model, progress management, information search, etc. These will be explained in detail below.

[0697] Hardware used

[0698] The present invention uses the following hardware:

[0699] Microphone: Collects voice input from the user.

[0700] Device: A device that converts speech to text and communicates with the server. This can be a smartphone, tablet, or computer.

[0701] Server: Runs the generative AI model, analyzes data, and manages progress.

[0702] Software used

[0703] The following software is used in the present invention:

[0704] speech_recognition: A Python library for converting speech to text.

[0705] requests: A Python library for sending HTTP requests.

[0706] Generative AI models: Artificial intelligence techniques for analyzing text and generating summaries and suggestions.

[0707] TextToSpeech: A library for converting text to speech.

[0708] System Operation

[0709] 1. The user speaks a command into the device's microphone, for example, "Start the next shaft inspection."

[0710] 2. The device converts the voice input into text in real time and sends the text data to the server.

[0711] 3. The server uses a generative AI model to analyze the received text data and generate appropriate actions and information, such as a summary and suggestion such as "Initiate shaft inspection protocol."

[0712] 4. The generated actions and information are sent to the terminal and notified to the user by display or voice.

[0713] 5. When the user inputs their progress via voice or text, the device sends the data to the server, which updates the progress management database. The next action or alert is generated and notified to the user via the device.

[0714] Specific examples

[0715] Consider a factory operator who wants to perform a shaft inspection on a production line. The operator gives a voice command such as "Start the next shaft inspection." The device converts the voice to text and sends it to the server. The server uses a generative AI model to analyze the message and return it to the device: "Starting shaft inspection protocol." The operator proceeds with the work according to the protocol and reports the progress by voice. The server updates the progress management database and notifies the operator of the next task or alert.

[0716] Prompt Sentence Examples

[0717] Here are some examples of specific prompts:

[0718] "We will begin testing Shaft A next week."

[0719] "What are the technical specifications of the new product?"

[0720] This system allows factory operators to give voice instructions, which enable progress management and the rapid acquisition of related information.

[0721] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0722] Step 1:

[0723] The user inputs voice instructions through the microphone of the terminal, and the input voice data is sent to the terminal.

[0724] Step 2:

[0725] The device receives the voice data and converts it into text data in real time. Specifically, it uses the speech_recognition library to analyze the voice and output it as text data.

[0726] Step 3:

[0727] The text data generated by the terminal is sent to the server using an HTTP request.

[0728] Step 4:

[0729] The server inputs the received text data into a generative AI model for data analysis. The AI ​​model analyzes the instructions based on the prompt and generates a summary and proposal. For example, in response to the instruction "Start the next shaft inspection," it generates the summary "Start the shaft inspection protocol."

[0730] Step 5:

[0731] The server sends the generated summary and proposal to the terminal as an HTTP response.

[0732] Step 6:

[0733] The device receives the summary and suggestions from the server and displays them to the user, either by displaying the summary as text on the screen or by announcing it aloud using the TextToSpeech library.

[0734] Step 7:

[0735] The user inputs progress by voice or text, and the data is received by the terminal. For example, the user may report by voice that "the shaft inspection is complete."

[0736] Step 8:

[0737] The device converts the voice data into text in real time and sends the text data to the server, again using an HTTP request.

[0738] Step 9:

[0739] The server updates the progress management database based on the received text data. For example, it moves "Shaft Inspection" from the in-progress task list to the completed task list.

[0740] Step 10:

[0741] Based on the results of the updates to the progress management database, the server generates the next action or alert, based on pre-defined rules and conditions.

[0742] Step 11:

[0743] The server generates actions and alerts and sends them to the terminal as HTTP responses.

[0744] Step 12:

[0745] The device notifies the user of actions and alerts from the server, either by displaying them on the device screen or by voice using the TextToSpeech library.

[0746] Step 13:

[0747] When a user wants to search for information, the user inputs a query by voice or text, and the device receives the input. For example, the user may input a query such as "What are the technical specifications of a new product?"

[0748] Step 14:

[0749] The device converts the voice data into text and sends the text data to the server, again using an HTTP request.

[0750] Step 15:

[0751] The server searches the database based on the text data to retrieve relevant information, then analyzes the retrieved information using a generative AI model to generate search results in a format appropriate for the user.

[0752] Step 16:

[0753] The search results generated by the server are sent to the terminal as an HTTP response.

[0754] Step 17:

[0755] The device displays the search results from the server to the user, either as text on the screen or as audio using the TextToSpeech library.

[0756] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0757] Understood. Below is the "Form for carrying out the invention".

[0758] The sales support AI app of this invention is a system that converts user voice input into text in real time and uses that data to summarize information and make suggestions using generative AI. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and by adjusting the content of suggestions based on the user's emotional state, it provides more effective sales support.

[0759] Business negotiation support

[0760] During business negotiations, the user launches the app and uses voice input. The device converts the voice input into text in real time using voice recognition technology and sends the text data to the server. The server then uses generative AI to analyze the text data, summarizes key points, and suggests next steps. The server then uses an emotion engine to analyze the user's emotions from the voice data and adjusts the proposal content based on those emotions. The adjusted summary and proposal are then displayed to the user via the device.

[0761] For example, if a user is introducing a new product during a sales meeting and the customer asks, "How much does this product cost?", the device will recognize the voice, convert it into text data, and send it to the server. The server will analyze the text data and recognize the user's emotions through an emotion engine. If the user seems nervous, it will generate a more polite explanation and return it to the device. As a result, the user can respond, "The price of this product is XX yen, but we are running a discount campaign for a limited time."

[0762] Progress management support

[0763] To update progress, the user inputs progress information by voice or text. The device performs voice recognition or text analysis on this information and sends it to the server as text data. The server updates the progress management database and generates the next action or alert. The server also uses an emotion engine to analyze the user's emotions when entering progress information, and generates a faster alert if the information is urgent or important. The generated action or alert is then notified to the user via the device.

[0764] For example, if a user types, "I need to set up my next meeting with Client A," the device recognizes this, converts it into text data, and sends it to the server. The server updates the progress management database, analyzes the user's state of tension using an emotion engine, and then generates a high-priority alert and sends it to the device. As a result, the user receives a real-time notification that "A meeting with Client A has been set up."

[0765] Information search support

[0766] In information searches, users input the information they are looking for by voice or text. The device performs voice recognition or text analysis on the input and sends the text data to a server. The server receives the query and searches a database for relevant information. The emotion engine adjusts the content of the search results to take into account the user's emotions. For example, if the user is in a hurry, the device will display the search results in a concise summary format. The device then displays the results to the user and provides voice feedback as needed.

[0767] For example, if a user says, "Tell me the technical specifications of this product," the device converts this into text data and sends it to the server. The server searches the database for the relevant technical specifications, and if the emotion engine detects the user's impatience, it generates a concise and easy-to-understand technical specification and returns it to the device. As a result, the user can quickly and accurately tell the user, "The technical specifications of this product are ____, and for more details, please see this link."

[0768] In this way, by incorporating an emotion engine, it is possible to respond according to the user's emotional state, further improving the quality and efficiency of sales activities.

[0769] The processing flow will be explained below.

[0770] Understood. Below, I will explain in detail the processing steps of the sales support AI app that combines an emotion engine.

[0771] Business negotiation support

[0772] Step 1:

[0773] The user launches the app by tapping the app icon and begins voice input.

[0774] Step 2:

[0775] The device activates the voice input function and records the user's voice in real time, which is then sent to a voice recognition algorithm.

[0776] Step 3:

[0777] The device recognizes the recorded voice and converts it into text data, which is then temporarily saved.

[0778] Step 4:

[0779] The device sends text and voice data to the server in real time, where it is uploaded via a network connection.

[0780] Step 5:

[0781] The server receives the text data and uses generative AI to summarize the main points, while simultaneously analyzing the audio data with an emotion engine to recognize the user's emotional state.

[0782] Step 6:

[0783] Based on the analysis, the server generates suggestions that match the user's emotions, for example, if the user is nervous, a suggestion with more polite explanations will be generated.

[0784] Step 7:

[0785] The server then sends the generated summary data and sentiment-based suggestions to the device, where the data is transferred over the network.

[0786] Step 8:

[0787] The device displays summary data and suggestions to the user, providing on-screen text and, optionally, audio feedback.

[0788] Progress management support

[0789] Step 1:

[0790] The user enters progress information into the app, using voice or text input functionality to enter progress details.

[0791] Step 2:

[0792] The device converts voice into text in real time and temporarily stores the input text data.

[0793] Step 3:

[0794] The device sends the text data to the server, where it is uploaded over a network connection.

[0795] Step 4:

[0796] The server receives the progress data and updates the progress management database, and the new progress information is recorded in the database.

[0797] Step 5:

[0798] The server generates the next action or alert based on the progress data, using an emotion engine to analyze the user's emotional state and evaluate the urgency and importance of the action or alert.

[0799] Step 6:

[0800] The server sends the generated actions and alerts to the terminal, and the data is transferred over the network.

[0801] Step 7:

[0802] The device notifies the user of the generated actions and alerts by displaying them on the screen and / or providing audio feedback.

[0803] Information search support

[0804] Step 1:

[0805] The user enters the information they want to search for by voice or text, taps the search button, and enters search keywords.

[0806] Step 2:

[0807] The device converts voice input into text in real time and temporarily stores the text data.

[0808] Step 3:

[0809] The device sends text data to the server, where it is uploaded via the network.

[0810] Step 4:

[0811] The server searches the database based on the received query and extracts data that matches the specified keywords.

[0812] Step 5:

[0813] The server generates search results and analyzes the user's emotional state using an emotion engine, and adjusts the display format of the search results based on the user's emotion.

[0814] Step 6:

[0815] The server sends the adjusted search results to the device, and the data is transferred over the network.

[0816] Step 7:

[0817] The device displays the search results to the user, providing text on the screen and, optionally, audio feedback.

[0818] By following these steps, a sales support AI app incorporating an emotion engine will provide responses that correspond to the user's emotional state, improving the quality and efficiency of sales activities.

[0819] Example 2

[0820] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0821] In the processing of voice input in sales activities, the challenge is to generate appropriate information summaries and suggestions that take the user's emotions into account, thereby improving the efficiency of progress management and information retrieval. Conventional systems can only respond uniformly, ignoring the user's emotions, and it is difficult to respond flexibly according to the user's situation.

[0822] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0823] In this invention, the server includes means for analyzing the user's emotions from voice data and adjusting the content of proposals based on the emotions, means for analyzing the user's emotions when entering progress and generating a prompt alert if the urgency or importance is high, and means for adjusting the display content in consideration of the user's emotional state when generating search results. This enables flexible responses according to the user's emotional state, improving the quality and efficiency of sales activities.

[0824] Understood. Below are definitions of important words.

[0825] "Voice input" refers to the process by which a user inputs voice through a microphone.

[0826] "Convert to text" refers to the process of converting audio data into written information in real time.

[0827] "Generative AI" refers to artificial intelligence technology that learns from large amounts of data and generates new information and suggestions.

[0828] "Emotion analysis" refers to the process of analyzing a user's emotional state from input voice and text data.

[0829] "Emotion engine" refers to software or hardware used to recognize the emotional state of a user.

[0830] "Adjusting suggestions" refers to the process of changing the content of generated suggestions depending on the user's emotional state.

[0831] "Progress information" refers to data that indicates the progress or task status related to a project or sales activity.

[0832] "Progress management database" refers to a database for managing and storing progress information.

[0833] "Generating an alert" refers to the process of creating a notification to draw attention to an event of high urgency or importance.

[0834] "Search query" means a command or question entered by a user into a system in search of specific information.

[0835] "Search Results" refers to relevant information retrieved from a database based on a user's search query.

[0836] "Adjusting display content" refers to the process of changing the format and detail of the information displayed to take into account the user's emotional state.

[0837] MODE FOR CARRYING OUT THE INVENTION

[0838] The sales support system of the present invention converts a user's voice input into text in real time, and uses generative AI to summarize the information and make suggestions based on that data. It also incorporates an emotion engine that recognizes the user's emotions, and can adjust the content of suggestions based on the user's emotional state. A specific embodiment of this system is described below.

[0839] Hardware and software used

[0840] This system uses the following hardware and software:

[0841] Devices (e.g. smartphones, tablets, laptops)

[0842] Server (cloud server or on-premise server)

[0843] Voice recognition technology (e.g., Google Cloud Speech-to-Text)

[0844] Generative AI models (e.g., OpenAI's GPT-3)

[0845] Sentiment analysis engine (e.g., Microsoft Azure's Text Analytics API)

[0846] Specific Example of the System

[0847] Business negotiation support

[0848] 1. The user launches the app and uses voice input, for example, "What is the price of the new product?"

[0849] 2. The device converts the voice input into text in real time using Google Cloud Speech-to-Text technology and sends the text data to the server.

[0850] 3. The server analyzes the received text data using a generative AI model (GPT-3), extracts key points, and suggests the next action. For example, it generates a suggestion such as, "The price of this product is XX yen."

[0851] 4. The server uses an emotion analysis engine to analyze the user's emotions from the voice data, and if the user is nervous, adjusts the explanation or suggestions to be more friendly.

[0852] 5. The adjusted summary and suggestions are displayed to the user via the device.

[0853] Example prompt sentence:

[0854] "New product pricing information that can be used in business negotiations"

[0855] Progress management support

[0856] 1. The user inputs status information by voice or text, for example, "I need to schedule my next meeting with Client A."

[0857] 2. The device uses voice recognition technology to convert the voice into text and sends the data to the server.

[0858] 3. The server updates the progress management database and generates the next action or alert, for example, "A meeting has been scheduled with Client A."

[0859] 4. The server uses an emotion analysis engine to analyze the user's emotions when entering progress, and if it determines that the situation is urgent, it generates a prompt alert.

[0860] 5. The device notifies the user of the generated action or alert.

[0861] Example prompt sentence:

[0862] "Schedule my next meeting"

[0863] Information search support

[0864] 1. The user enters the information they are looking for by voice or text, for example, "What are the technical specifications for this product?"

[0865] 2. The device uses voice recognition technology to convert the voice into text and sends the data to the server.

[0866] 3. The server searches the database for relevant information and generates search results, such as "The technical specifications of this product are ____."

[0867] 4. When generating search results, the server analyzes the user's emotional state using an emotion analysis engine and adjusts the displayed content accordingly. For example, if the user is anxious, the server will provide information in a concise format.

[0868] 5. The device displays the adjusted search results to the user and provides audio feedback if necessary.

[0869] Example prompt sentence:

[0870] "What are the technical specifications of the product?"

[0871] This system aims to improve the quality and efficiency of sales activities by enabling flexible responses according to the user's emotional state.

[0872] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0873] Sales Support Processing Steps

[0874] Step 1:

[0875] The user launches the app and performs voice input.

[0876] Specifically, the user opens the app on their smartphone and asks, "What is the price of the new product?"

[0877] Input: User's voice data

[0878] Output: Audio input data

[0879] Step 2:

[0880] The device converts voice input into text in real time using speech recognition technology (e.g., Google Cloud Speech-to-Text).

[0881] Specifically, the device's microphone captures audio and converts it into text via a cloud service.

[0882] Input: Voice input data

[0883] Output: Text data

[0884] Step 3:

[0885] The terminal transmits the converted text data to the server.

[0886] Specifically, the device sends text data to the server over a secure HTTPS connection.

[0887] Input: Text data

[0888] Output: The text data sent

[0889] Step 4:

[0890] The server uses a generative AI model (e.g., GPT-3) to analyze the text data, extract key points, and suggest the next action.

[0891] Specifically, the server passes the text data to the analysis engine, and the generative AI model generates information about the "price of the new product."

[0892] Input: Text data sent

[0893] Output: Summary of proposal

[0894] Step 5:

[0895] The server uses a sentiment analysis engine (for example, Microsoft Azure's Text Analytics API) to analyze the user's emotions from the voice data and adjusts the suggestions based on those emotions.

[0896] Specifically, the server transfers voice data to an analysis engine, and if it detects that the user is in a tense state, it changes the suggestions to be more polite.

[0897] Input: Audio data, summarized proposal

[0898] Output: Sentiment-adjusted recommendations

[0899] Step 6:

[0900] The server sends the adjusted summary and suggestions to the terminal.

[0901] Specifically, the adjusted proposal is sent to the device over a secure connection.

[0902] Input: Emotion-adjusted recommendations

[0903] Output: Adjustment proposals submitted

[0904] Step 7:

[0905] The device displays the adjusted proposal to the user.

[0906] Specifically, the device will display on the screen, "The price of this product is XX yen. We are running a discount campaign for a limited time."

[0907] Input: Adjustment proposal submitted

[0908] Output: The suggestions displayed to the user

[0909] Progress management support processing steps

[0910] Step 1:

[0911] The user enters progress information by voice or text.

[0912] Specifically, the user might say or type in text, "I need to schedule my next meeting with Client A."

[0913] Input: Audio or text data

[0914] Output: Progress input data

[0915] Step 2:

[0916] The device converts the voice input into text using voice recognition technology and sends the data to the server.

[0917] Specifically, the terminal converts the progress input into text and securely transmits it to the server.

[0918] Input: Progress input data

[0919] Output: Progress data sent

[0920] Step 3:

[0921] The server updates the progress management database and generates the next action or alert.

[0922] As a specific operation, the server adds new progress information to the database and generates an action as "meeting set up."

[0923] Input: Submitted progress data

[0924] Output: Updated progress data, generated actions

[0925] Step 4:

[0926] The server uses an emotion analysis engine to analyze the user's emotions when entering progress information, and generates a prompt alert if the information is of high urgency or importance.

[0927] Specifically, the server analyzes the emotions expressed when entering progress, and if it determines that the situation is urgent, it generates a high-priority alert.

[0928] Input: Progress data, user emotion data

[0929] Output: The generated alert

[0930] Step 5:

[0931] The terminal notifies the user of the generated actions and alerts.

[0932] Specifically, the device will display a notification on the screen saying, "A meeting with Client A has been set up."

[0933] Input: Generated action, alert

[0934] Output: Information reported to the user

[0935] Information retrieval support processing steps

[0936] Step 1:

[0937] The user inputs the information they are looking for by voice or text.

[0938] Specifically, the user says, "Tell me the technical specifications of this product."

[0939] Input: Audio or text data

[0940] Output: Search query

[0941] Step 2:

[0942] The device converts the input into text using voice recognition technology and sends the data to the server.

[0943] Specifically, the device converts the voice input into text and sends it to the server.

[0944] Input: search query

[0945] Output: Submitted search query data

[0946] Step 3:

[0947] The server receives the query and retrieves the relevant information from a database.

[0948] Specifically, the server queries a product database to obtain the relevant technical specification information.

[0949] Input: Submitted search query data

[0950] Output: Related information

[0951] Step 4:

[0952] When generating search results, the server analyzes the user's emotional state using an emotion analysis engine and adjusts the display content.

[0953] Specifically, if the server determines that the user is in a hurry, it displays the technical specifications in a concise and easy-to-understand format.

[0954] Input: Related information, user emotion data

[0955] Output: Search results tailored based on sentiment

[0956] Step 5:

[0957] The device displays the adjusted search results to the user and provides audio feedback if necessary.

[0958] Specifically, the device will display on the screen, "The technical specifications of this product are XX. Please refer to this link for details," and the voice assistant will also read it out loud.

[0959] Input: Refined search results

[0960] Output: Search results displayed to the user

[0961] (Application example 2)

[0962] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0963] Sales activities in brick-and-mortar stores require immediate responses to customer questions and reactions, but conventional systems have difficulty taking customer emotions into account, often failing to provide appropriate information or suggestions. Furthermore, in situations where salespeople are at a loss as to how to respond, a system is needed that can generate and provide appropriate suggestions in real time. The present invention aims to solve these problems and provide a system that supports salespeople in effectively responding to customers.

[0964] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for using an emotion analysis engine that analyzes emotions from audio data and video data, means for generating summaries and proposals using generative AI, and means for the server to transmit summaries and proposals adjusted based on the emotion analysis. This makes it possible to provide information and proposals in real time that take customer emotions into consideration.

[0965] "User" refers to an individual or entity who uses the System to provide voice or text input.

[0966] A "terminal" is a device operated by a user that has the function of sending voice input and text data to a server.

[0967] "Server" refers to a central processing unit that receives text data sent from a device, processes the information using generative AI or an emotion analysis engine, and returns the results to the device.

[0968] "Generative AI" refers to an artificial intelligence engine that analyzes text data and automatically generates summaries of information and suggestions.

[0969] An "emotion analysis engine" is an engine that has the ability to analyze a user's emotions from audio and video data and adjust the content of suggestions based on that information.

[0970] "Voice input" refers to user voice data collected by the terminal, which is converted into text in real time.

[0971] "Text data" refers to information in which voice input is converted into text in real time.

[0972] A "summary" refers to information that is concisely summarized by generative AI that analyzes text data and extracts important points.

[0973] "Suggestion" refers to the next action or information the user should take, generated by generative AI based on text data.

[0974] "Progress management database" refers to a database for storing and managing user progress information.

[0975] An "alert" refers to urgent or important information that is notified to the user when the progress management database is updated.

[0976] "Search results" refers to information that the server searches for related information from the database and presents to the user.

[0977] The present invention is a system that supports sales activities in brick-and-mortar stores, and makes it possible to provide information and suggestions in real time based on the user's voice input and video data.

[0978] System Overview

[0979] The system includes the following main elements:

[0980] 1. User (salesperson): Operates the system and provides voice input.

[0981] 2. Terminal: A device operated by the user (e.g., smart glasses) that collects voice input and sends it to a server.

[0982] 3. Server: Analyzes data sent from the device and processes the information using generative AI and an emotion analysis engine.

[0983] Hardware and Software Details

[0984] Smart glasses: Devices with AR capabilities and microphones (e.g., Google Glass) that collect audio and visual data from the user.

[0985] Speech Recognition API: Uses the Google Cloud Speech-to-Text API to convert voice input to text in real time.

[0986] Generative AI model: Uses OpenAI GPT to analyze text data and generate information summaries and suggestions.

[0987] Sentiment analysis engine: Uses the Microsoft Azure Emotion API to analyze emotions from audio and video data.

[0988] Cloud server: Data processing is performed using Amazon Web Services (AWS).

[0989] Operation explanation

[0990] Receiving audio input

[0991] A user (salesperson) puts on the smart glasses during a sales meeting and starts a conversation with a customer. The microphone in the smart glasses collects the salesperson's voice input and temporarily stores the voice data. The stored voice data is converted into text data in real time by the Google Cloud Speech-to-Text API.

[0992] Sending and analyzing text data

[0993] The converted text data is sent to a cloud server (AWS). A generative AI model (OpenAI GPT) on the cloud server analyzes the text data, summarizes key points, and generates next actions and suggestions. In parallel, the audio and video data is analyzed by the Microsoft Azure Emotion API to determine the customer's emotions. Based on the results of the emotion analysis, the suggestions generated by the generative AI are adjusted appropriately.

[0994] Viewing Proposals

[0995] The final tailored summary and recommendations are then displayed on the smart glasses display, allowing the salesperson to respond appropriately to the customer.

[0996] Specific examples

[0997] For example, if a customer asks a salesperson, "How much does this product cost?", the smart glasses will collect the voice and convert it into text data in real time. This text data is sent to a server and analyzed by a generative AI model. The sentiment analysis engine will analyze the customer's emotions, and if the customer seems impatient, a short, concise, and unassuming answer will be generated. As a result, the salesperson can respond, "The price of this product is XX yen, but we are running a discount campaign for a limited time."

[0998] Prompt Sentence Examples

[0999] Convert user questions into text in real time, analyze customer sentiment, and generate optimal responses.

[1000] Example input:

[1001] Salesperson: "Look at this product, it has special features."

[1002] Customer: "What's the price?"

[1003] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1004] Step 1:

[1005] A user (salesperson) puts on smart glasses and starts a sales negotiation. As input, the microphone in the smart glasses collects voice data of the conversation with the customer. As output, the voice data is temporarily stored in the built-in memory. Specifically, the smart glasses continuously capture voice through the microphone.

[1006] Step 2:

[1007] The collected voice data is converted into text data in real time. As input, the voice data collected in step 1 is sent to the Google Cloud Speech-to-Text API. As output, the voice data is converted into text data. Specifically, the device (smart glasses) calls the Google Cloud Speech-to-Text API to convert the voice data into text.

[1008] Step 3:

[1009] The converted text data is sent to the server. As input, the text data converted in step 2 is sent to the cloud server (AWS). As output, the text data is saved on the server. In concrete terms, the device sends the text data to the cloud server via the Internet.

[1010] Step 4:

[1011] The server analyzes the text data and generates a summary and suggestions based on key points. The text data sent to the server in step 3 is received as input. The summary and suggestions are generated as output. Specifically, a generative AI model (OpenAI GPT) on the server analyzes the text data, extracts key elements, and generates suggestions.

[1012] Step 5:

[1013] In parallel, the audio and video data are sent to the emotion analysis engine on the server for analysis. As input, the audio data collected in step 1 and the video data captured by the smart glasses camera are sent to the server. As output, customer emotion data is generated. Specifically, the server uses the Microsoft Azure Emotion API to analyze emotions from the audio and video.

[1014] Step 6:

[1015] The generated proposals are adjusted based on the results of the sentiment analysis engine. The summaries and proposals generated in step 4 and the emotion data generated in step 5 are used as input. The output is a summary and proposal adjusted based on emotion. Specifically, the generative AI model on the server uses the emotion data to optimize the summaries and proposals.

[1016] Step 7:

[1017] The final summary and suggestions are sent to the terminal and displayed to the user. As input, the adjusted summary and suggestions generated in step 6 are sent from the cloud server to the terminal. As output, the suggestions are displayed on the display of the smart glasses. Specifically, the server sends the summary and suggestions to the smart glasses via the Internet, and the smart glasses display them.

[1018] Prompt Sentence Examples

[1019] Convert user questions into text in real time, analyze customer sentiment, and generate optimal responses.

[1020] Example input:

[1021] Salesperson: "Look at this product, it has special features."

[1022] Customer: "What's the price?"

[1023] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1024] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1025] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1026] [Third embodiment]

[1027] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1028] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1030] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1031] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1032] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1034] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1035] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1037] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1038] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1039] Understood. Below is the "Form for carrying out the invention".

[1040] In one embodiment of the invention, the system begins with a user providing instructions through voice input, which the device converts in real time into text and sends to the server. The server then uses generative AI to analyze the text data, generating summaries and suggestions, which are then sent back to the device for display to the user. When the user inputs progress, this is recorded in a progress management database, and the user is notified of next actions or alerts. Furthermore, when the user searches for specific information, the server searches for related data and provides the results to the user.

[1041] Business negotiation support

[1042] When a user launches the app, the app activates its voice input function. The device converts the user's speech into text in real time and sends the text data to a server. The server then uses generative AI to analyze the text data, summarize key points, and suggest next steps. These summaries and suggestions are displayed to the user on the device.

[1043] For example, suppose a user is introducing a new product during a business meeting and the client asks, "What is the price of this product?" In this case, the device recognizes the voice as a "price question" and sends the text data to the server. The server analyzes past data, generates appropriate price information, and returns it to the device. As a result, the user can respond to the client by saying, "The price of this product is XX yen."

[1044] Progress management support

[1045] In progress management, the user inputs the progress of each business negotiation or task by voice or text. The device recognizes this and sends the data to the server. The server updates the progress management database and generates the next action or alert. These actions and alerts are notified to the user via the device.

[1046] For example, when a user says, "Set up a meeting with Client A next week," the device converts the voice to text and sends it to the server. The server updates the progress management database and generates an alert to set up the next meeting. The device notifies the user when the meeting date approaches.

[1047] Information search support

[1048] With the information search function, users search for the information they need using voice or text. The device analyzes this input and sends the query to the server. The server then searches for relevant information in a database, generates results, and sends them to the device. The user can then view the search results through the device.

[1049] For example, if a user says, "I want to know the technical specifications of a new product," the device converts the instruction into text and sends it to the server. The server then searches the database for the new product's technical specifications and returns them to the device. The user can then view the specifications on the screen.

[1050] As a result, the system standardizes the quality of business negotiations, improves the efficiency of progress management, and enables quick and accurate information acquisition.

[1051] The processing flow will be explained below.

[1052] Understood. Below are the specific processing steps.

[1053] Business negotiation support

[1054] Step 1:

[1055] The user launches the app by tapping the app icon.

[1056] Step 2:

[1057] The terminal activates the voice input function and starts recording the user's voice in real time.

[1058] Step 3:

[1059] The device converts the recorded audio into text data in real time, using a speech recognition algorithm to output the audio as text.

[1060] Step 4:

[1061] The terminal transmits the text data to the server, and the text data is uploaded to the server via the network.

[1062] Step 5:

[1063] The server receives the text data and uses generative AI to summarize the information, including important points and next actions.

[1064] Step 6:

[1065] The server generates information summaries and suggestions, and creates optimal suggestions based on the results of analyzing the text data.

[1066] Step 7:

[1067] The server transmits the generated summary data and the proposal to the terminal, and transmits the data to the terminal via the network.

[1068] Step 8:

[1069] The device displays summary data and suggestions to the user, either as text on the screen or as audio feedback.

[1070] Progress management support

[1071] Step 1:

[1072] The user types a progress update into the app or gives a voice command. They tap the progress update button and enter the information.

[1073] Step 2:

[1074] The terminal performs voice recognition or text analysis on the input content to generate text data.

[1075] Step 3:

[1076] The device sends the text data to the server, where it is uploaded via the network.

[1077] Step 4:

[1078] The server receives the progress data and updates the progress management database, recording the new progress information.

[1079] Step 5:

[1080] The server analyzes the progress and generates the next action or alert: Create the necessary tasks or alerts.

[1081] Step 6:

[1082] The server sends the generated tasks and alerts to the terminal and returns the data via the network.

[1083] Step 7:

[1084] The device notifies the user of tasks and alerts, displaying them on the screen and providing audio feedback where appropriate.

[1085] Information search support

[1086] Step 1:

[1087] When a user wants to search for specific information, they input it by voice or text, then tap the search button.

[1088] Step 2:

[1089] The terminal performs voice recognition or text analysis on the input content to generate text data.

[1090] Step 3:

[1091] The device sends the text data to the server, where it is uploaded via the network.

[1092] Step 4:

[1093] The server receives the query and searches the database to extract information that matches the specified keywords.

[1094] Step 5:

[1095] The server generates search results and sends them to the device, organizes the results, and sends the data to the device.

[1096] Step 6:

[1097] The device displays the search results to the user, displaying them on the screen and providing audio feedback if necessary.

[1098] By following the above steps, this system provides functions such as support for business negotiations, progress management, and information search.

[1099] Example 1

[1100] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1101] Conventional sales negotiation support systems, progress management systems, and information search systems each function independently, resulting in low user convenience. Furthermore, there is a lack of systems that can efficiently convert voice input into text and then centrally manage the subsequent analysis, proposals, progress management, and information search. This creates challenges for users, making it difficult to standardize the quality of sales negotiations, efficiently manage progress, and quickly obtain information.

[1102] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1103] In this invention, the server includes means for analyzing information using a generative artificial intelligence model and generating summaries and proposals, means for updating the progress management database and generating next actions and alerts, and means for searching the database for related information, generating search results, and sending them to the terminal. This allows for centralized management of business negotiation support, progress management, and information search, enabling users to efficiently give instructions and obtain information.

[1104] "User" refers to any individual or corporation that uses this system.

[1105] "Voice input" refers to a means by which a user gives instructions to a system using voice.

[1106] "Terminal" refers to a hardware device that converts speech to text, transmits the data to a server, and displays the results.

[1107] "Server" refers to a computer system that uses generative artificial intelligence models to analyze text data and generate summaries, suggestions, progress management, and information search results.

[1108] A "generative artificial intelligence model" refers to an algorithm or program that analyzes input text data and generates summaries or suggestions based on it.

[1109] "Text data" refers to data resulting from converting voice input into text.

[1110] "Analyzing information" refers to the process of understanding input text data and generating summaries or suggestions based on that content.

[1111] A "summary" refers to a concise summary of the important points of text data.

[1112] "Suggestion" refers to showing the user the next action or countermeasure to be taken based on the analysis results.

[1113] A "progress management database" refers to a database for recording and managing the progress of business negotiations and tasks.

[1114] "Action" refers to the specific next steps or actions to be taken.

[1115] "Alert" refers to a warning or reminder that notifies the user of important information or progress.

[1116] "Information search" refers to a search performed by a user to obtain specific information.

[1117] "Database" refers to a data storage system in which related information is stored.

[1118] "Search Results" refers to a server-generated answer or collection of data in response to an information search query.

[1119] This invention is a system that unifies the management of business negotiation support, progress management, and information search. This system is composed of three main components: users, terminals, and servers.

[1120] Business negotiation support

[1121] When a user launches the app, the device activates the voice input function. When the user speaks, the device converts the speech into text in real time and sends the text data to the server. The server then analyzes the text data using a generative artificial intelligence model (e.g., OpenAI's GPT series), summarizes the key points, and suggests the next action to take. The suggested summary and action are displayed to the user on the device.

[1122] For example, if a user is introducing a new product and the client asks, "What is the price of this product?", the device converts the speech into text "Question about price" and sends it to the server. The server analyzes past data, generates appropriate price information, and returns it to the device. The user can then reply to the client, "The price of this product is XX yen."

[1123] Progress management

[1124] When the user inputs progress by voice or text, the device sends the data to the server, which updates the progress management database (e.g., PostgreSQL) and generates the next action or alert. The generated action or alert is then notified to the user via the device.

[1125] For example, if a user says, "Set up a meeting with Client A next week," the device converts the speech into text and sends it to the server. The server updates the progress management database and generates an alert to set up the next meeting. When the meeting date approaches, the device notifies the user.

[1126] Information Search

[1127] When a user searches for the information they need by voice or text, the device sends the query to the server, which then searches for relevant information in a database (e.g., Elasticsearch), generates results, and sends them to the device. The user can then view the search results through the device.

[1128] For example, if a user says, "I want to know the technical specifications of a new product," the device converts the instruction into text and sends it to the server. The server then searches the database for the new product's technical specifications and returns them to the device. The user can then view the specifications on the device's screen.

[1129] As a result, this system can standardize the quality of business negotiations, improve the efficiency of progress management, and enable quick and accurate information acquisition.

[1130] Prompt Sentence Examples

[1131] For sales support: "Please summarize the notes from your meeting with Client X and tell me what action to take next."

[1132] For progress management support: "How is project Y progressing? What's the next action?"

[1133] For information search assistance: "Please tell me the technical specifications for new product Z."

[1134] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1135] Understood. Below we will explain in detail the processing flow of the system program, divided into processing steps.

[1136] Business negotiation support

[1137] Step 1:

[1138] A user launches an app and activates the voice input function. The input is the command to launch the app, and the output is that voice input is enabled.

[1139] Step 2:

[1140] The user provides voice input, which the device converts into text in real time. The input is the user's voice data, and the output is text data. Specifically, the device uses voice recognition software (e.g., Google Speech-to-Text API).

[1141] Step 3:

[1142] The terminal sends the generated text data to the server. The input is the text data, and the output is the data sent to the server. Specifically, the terminal sends data to the server using an HTTPS request.

[1143] Step 4:

[1144] The server uses a generative AI model to analyze the transmitted text data. The input is text data, and the output is a summary and proposed data. Specifically, the server uses a generative AI model such as GPT-4 to perform the analysis and generation process.

[1145] Step 5:

[1146] The server sends the generated summary and proposal to the terminal. The input is the summary and proposal data, and the output is the data to be sent to the terminal. Specifically, the server sends the generated results in JSON format.

[1147] Step 6:

[1148] The terminal displays the received summary and suggestions to the user. The input is the data received from the server, and the output is the visualization to the user. Specifically, the terminal displays the data using a GUI.

[1149] Progress management

[1150] Step 1:

[1151] The user inputs their progress by voice or text. The input is the user's voice or text data, and the output is text data. Specifically, the device recognizes the voice and converts it into text.

[1152] Step 2:

[1153] The terminal sends the text data to the server. The input is the text data, and the output is the data sent to the server. Specific operations use an HTTPS request.

[1154] Step 3:

[1155] The server updates the progress management database. The input is text data, and the output is the updated database. Specifically, the server updates the database using an SQL query.

[1156] Step 4:

[1157] The server generates the next action or alert. The input is the updated database information, and the output is the action / alert data. Specifically, the server checks for pending tasks and generates the next action or alert.

[1158] Step 5:

[1159] The device notifies the user of the generated actions and alerts. The input is the action / alert data sent from the server, and the output is the notification to the user. Specifically, the device performs push notifications and in-app notifications.

[1160] Information Search

[1161] Step 1:

[1162] The user searches for the information they need by voice or text. The input is the user's query (voice or text), and the output is the query data. Specifically, the device performs voice recognition and converts it into text.

[1163] Step 2:

[1164] The device sends the query to the server. The input is the query data, and the output is the data sent to the server. Specific operations use an HTTPS request.

[1165] Step 3:

[1166] The server searches for relevant information from a database. The input is the query data, and the output is the search result data. Specifically, the server uses a search engine such as Elasticsearch.

[1167] Step 4:

[1168] The server generates search results and sends them to the terminal. The input is the search result data, and the output is the data to be sent to the terminal. Specifically, the results are sent in JSON format.

[1169] Step 5:

[1170] The terminal displays the search results to the user. The input is the data received from the server, and the output is the display to the user. Specifically, the terminal displays the data using a GUI.

[1171] (Application example 1)

[1172] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1173] Modern manufacturing factory operations require automation and real-time progress management. In particular, there is a lack of technology that allows factory operators to give work instructions to robots via voice and efficiently manage their progress, resulting in a decline in production efficiency and work accuracy. There is also a need for a system that can quickly search for and provide necessary information.

[1174] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1175] In this invention, the server includes: means for a user to give instructions through voice input and for the terminal to convert the voice into text in real time; means for sending text data to the server and for the server to generate a summary and proposal of the information using a generative AI model; means for the terminal to display the summary and proposal sent from the server to the user; means for the terminal to convert the voice into text and then send the text to the server; means for the server to generate appropriate actions and information in response to the instructions and send the results to the terminal; means for the user to input progress and tasks by voice or text and for the terminal to send them to the server; means for the server to update the progress management database and generate the next action or alert; and means for the terminal to notify the user of the generated action or alert. This enables factory operators to give instructions to robots by voice and quickly manage progress and obtain related information based on those instructions.

[1176] A "user" is a person who uses the system to give instructions and input progress.

[1177] "Voice input" is a method in which a user gives instructions or information to a terminal by voice.

[1178] A "terminal" is a device that converts voice input from a user into text and transmits the text data to a server.

[1179] "Real-time" is a time characteristic that means processing data and returning results almost instantaneously.

[1180] "Means for converting to text" refers to technology or devices for converting voice data into character data.

[1181] "Text data" is a data format in which voice is converted into text.

[1182] A "server" is a central computer system that receives text data and performs analysis and information generation.

[1183] A "generative AI model" is an algorithm that uses artificial intelligence technology to analyze received data and generate information summaries and suggestions.

[1184] "Means of summarizing and suggesting information" is the process of using a generative AI model to extract important information from the data received and present the user with the next action to take.

[1185] The "means for displaying the summary and suggestions" refers to a technique by which the terminal presents the summary and suggestions sent from the server to the user visually or audibly.

[1186] "Progress" is data that indicates the progress of a particular task or project.

[1187] The "progress management database" is a database system for storing and managing progress information entered by the user.

[1188] "Means for generating actions and alerts" refers to the process for generating next actions and alerts based on the progress management database.

[1189] "Means for notifying actions and alerts" refers to techniques for notifying users of generated actions and alerts.

[1190] "Means for searching information" refers to the technology that allows a user to send information searched for by voice or text to a server and return the results.

[1191] The "means for displaying search results" refers to a technique for visually or audibly presenting the search results sent from the server to the user.

[1192] The system required to implement this invention includes functions such as voice input, real-time text conversion, generative AI model, progress management, information search, etc. These will be explained in detail below.

[1193] Hardware used

[1194] The present invention uses the following hardware:

[1195] Microphone: Collects voice input from the user.

[1196] Device: A device that converts speech to text and communicates with the server. This can be a smartphone, tablet, or computer.

[1197] Server: Runs the generative AI model, analyzes data, and manages progress.

[1198] Software used

[1199] The following software is used in the present invention:

[1200] speech_recognition: A Python library for converting speech to text.

[1201] requests: A Python library for sending HTTP requests.

[1202] Generative AI models: Artificial intelligence techniques for analyzing text and generating summaries and suggestions.

[1203] TextToSpeech: A library for converting text to speech.

[1204] System Operation

[1205] 1. The user speaks a command into the device's microphone, for example, "Start the next shaft inspection."

[1206] 2. The device converts the voice input into text in real time and sends the text data to the server.

[1207] 3. The server uses a generative AI model to analyze the received text data and generate appropriate actions and information, such as a summary and suggestion such as "Initiate shaft inspection protocol."

[1208] 4. The generated actions and information are sent to the terminal and notified to the user by display or voice.

[1209] 5. When the user inputs their progress via voice or text, the device sends the data to the server, which updates the progress management database. The next action or alert is generated and notified to the user via the device.

[1210] Specific examples

[1211] Consider a factory operator who wants to perform a shaft inspection on a production line. The operator gives a voice command such as "Start the next shaft inspection." The device converts the voice to text and sends it to the server. The server uses a generative AI model to analyze the message and return it to the device: "Starting shaft inspection protocol." The operator proceeds with the work according to the protocol and reports the progress by voice. The server updates the progress management database and notifies the operator of the next task or alert.

[1212] Prompt Sentence Examples

[1213] Here are some examples of specific prompts:

[1214] "We will begin testing Shaft A next week."

[1215] "What are the technical specifications of the new product?"

[1216] This system allows factory operators to give voice instructions, which enable progress management and the rapid acquisition of related information.

[1217] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1218] Step 1:

[1219] The user inputs voice instructions through the microphone of the terminal, and the input voice data is sent to the terminal.

[1220] Step 2:

[1221] The device receives the voice data and converts it into text data in real time. Specifically, it uses the speech_recognition library to analyze the voice and output it as text data.

[1222] Step 3:

[1223] The text data generated by the terminal is sent to the server using an HTTP request.

[1224] Step 4:

[1225] The server inputs the received text data into a generative AI model for data analysis. The AI ​​model analyzes the instructions based on the prompt and generates a summary and proposal. For example, in response to the instruction "Start the next shaft inspection," it generates the summary "Start the shaft inspection protocol."

[1226] Step 5:

[1227] The server sends the generated summary and proposal to the terminal as an HTTP response.

[1228] Step 6:

[1229] The device receives the summary and suggestions from the server and displays them to the user, either by displaying the summary as text on the screen or by announcing it aloud using the TextToSpeech library.

[1230] Step 7:

[1231] The user inputs progress by voice or text, and the data is received by the terminal. For example, the user may report by voice that "the shaft inspection is complete."

[1232] Step 8:

[1233] The device converts the voice data into text in real time and sends the text data to the server, again using an HTTP request.

[1234] Step 9:

[1235] The server updates the progress management database based on the received text data. For example, it moves "Shaft Inspection" from the in-progress task list to the completed task list.

[1236] Step 10:

[1237] Based on the results of the updates to the progress management database, the server generates the next action or alert, based on pre-defined rules and conditions.

[1238] Step 11:

[1239] The server generates actions and alerts and sends them to the terminal as HTTP responses.

[1240] Step 12:

[1241] The device notifies the user of actions and alerts from the server, either by displaying them on the device screen or by voice using the TextToSpeech library.

[1242] Step 13:

[1243] When a user wants to search for information, the user inputs a query by voice or text, and the device receives the input. For example, the user may input a query such as "What are the technical specifications of a new product?"

[1244] Step 14:

[1245] The device converts the voice data into text and sends the text data to the server, again using an HTTP request.

[1246] Step 15:

[1247] The server searches the database based on the text data to retrieve relevant information, then analyzes the retrieved information using a generative AI model to generate search results in a format appropriate for the user.

[1248] Step 16:

[1249] The search results generated by the server are sent to the terminal as an HTTP response.

[1250] Step 17:

[1251] The device displays the search results from the server to the user, either as text on the screen or as audio using the TextToSpeech library.

[1252] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1253] Understood. Below is the "Form for carrying out the invention".

[1254] The sales support AI app of this invention is a system that converts user voice input into text in real time and uses that data to summarize information and make suggestions using generative AI. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and by adjusting the content of suggestions based on the user's emotional state, it provides more effective sales support.

[1255] Business negotiation support

[1256] During business negotiations, the user launches the app and uses voice input. The device converts the voice input into text in real time using voice recognition technology and sends the text data to the server. The server then uses generative AI to analyze the text data, summarizes key points, and suggests next steps. The server then uses an emotion engine to analyze the user's emotions from the voice data and adjusts the proposal content based on those emotions. The adjusted summary and proposal are then displayed to the user via the device.

[1257] For example, if a user is introducing a new product during a sales meeting and the customer asks, "How much does this product cost?", the device will recognize the voice, convert it into text data, and send it to the server. The server will analyze the text data and recognize the user's emotions through an emotion engine. If the user seems nervous, it will generate a more polite explanation and return it to the device. As a result, the user can respond, "The price of this product is XX yen, but we are running a discount campaign for a limited time."

[1258] Progress management support

[1259] To update progress, the user inputs progress information by voice or text. The device performs voice recognition or text analysis on this information and sends it to the server as text data. The server updates the progress management database and generates the next action or alert. The server also uses an emotion engine to analyze the user's emotions when entering progress information, and generates a faster alert if the information is urgent or important. The generated action or alert is then notified to the user via the device.

[1260] For example, if a user types, "I need to set up my next meeting with Client A," the device recognizes this, converts it into text data, and sends it to the server. The server updates the progress management database, analyzes the user's state of tension using an emotion engine, and then generates a high-priority alert and sends it to the device. As a result, the user receives a real-time notification that "A meeting with Client A has been set up."

[1261] Information search support

[1262] In information searches, users input the information they are looking for by voice or text. The device performs voice recognition or text analysis on the input and sends the text data to a server. The server receives the query and searches a database for relevant information. The emotion engine adjusts the content of the search results to take into account the user's emotions. For example, if the user is in a hurry, the device will display the search results in a concise summary format. The device then displays the results to the user and provides voice feedback as needed.

[1263] For example, if a user says, "Tell me the technical specifications of this product," the device converts this into text data and sends it to the server. The server searches the database for the relevant technical specifications, and if the emotion engine detects the user's impatience, it generates a concise and easy-to-understand technical specification and returns it to the device. As a result, the user can quickly and accurately tell the user, "The technical specifications of this product are ____, and for more details, please see this link."

[1264] In this way, by incorporating an emotion engine, it is possible to respond according to the user's emotional state, further improving the quality and efficiency of sales activities.

[1265] The processing flow will be explained below.

[1266] Understood. Below, I will explain in detail the processing steps of the sales support AI app that combines an emotion engine.

[1267] Business negotiation support

[1268] Step 1:

[1269] The user launches the app by tapping the app icon and begins voice input.

[1270] Step 2:

[1271] The device activates the voice input function and records the user's voice in real time, which is then sent to a voice recognition algorithm.

[1272] Step 3:

[1273] The device recognizes the recorded voice and converts it into text data, which is then temporarily saved.

[1274] Step 4:

[1275] The device sends text and voice data to the server in real time, where it is uploaded via a network connection.

[1276] Step 5:

[1277] The server receives the text data and uses generative AI to summarize the main points, while simultaneously analyzing the audio data with an emotion engine to recognize the user's emotional state.

[1278] Step 6:

[1279] Based on the analysis, the server generates suggestions that match the user's emotions, for example, if the user is nervous, a suggestion with more polite explanations will be generated.

[1280] Step 7:

[1281] The server then sends the generated summary data and sentiment-based suggestions to the device, where the data is transferred over the network.

[1282] Step 8:

[1283] The device displays summary data and suggestions to the user, providing on-screen text and, optionally, audio feedback.

[1284] Progress management support

[1285] Step 1:

[1286] The user enters progress information into the app, using voice or text input functionality to enter progress details.

[1287] Step 2:

[1288] The device converts voice into text in real time and temporarily stores the input text data.

[1289] Step 3:

[1290] The device sends the text data to the server, where it is uploaded over a network connection.

[1291] Step 4:

[1292] The server receives the progress data and updates the progress management database, and the new progress information is recorded in the database.

[1293] Step 5:

[1294] The server generates the next action or alert based on the progress data, using an emotion engine to analyze the user's emotional state and evaluate the urgency and importance of the action or alert.

[1295] Step 6:

[1296] The server sends the generated actions and alerts to the terminal, and the data is transferred over the network.

[1297] Step 7:

[1298] The device notifies the user of the generated actions and alerts by displaying them on the screen and / or providing audio feedback.

[1299] Information search support

[1300] Step 1:

[1301] The user enters the information they want to search for by voice or text, taps the search button, and enters search keywords.

[1302] Step 2:

[1303] The device converts voice input into text in real time and temporarily stores the text data.

[1304] Step 3:

[1305] The device sends text data to the server, where it is uploaded via the network.

[1306] Step 4:

[1307] The server searches the database based on the received query and extracts data that matches the specified keywords.

[1308] Step 5:

[1309] The server generates search results and analyzes the user's emotional state using an emotion engine, and adjusts the display format of the search results based on the user's emotion.

[1310] Step 6:

[1311] The server sends the adjusted search results to the device, and the data is transferred over the network.

[1312] Step 7:

[1313] The device displays the search results to the user, providing text on the screen and, optionally, audio feedback.

[1314] By following these steps, a sales support AI app incorporating an emotion engine will provide responses that correspond to the user's emotional state, improving the quality and efficiency of sales activities.

[1315] Example 2

[1316] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1317] In the processing of voice input in sales activities, the challenge is to generate appropriate information summaries and suggestions that take the user's emotions into account, thereby improving the efficiency of progress management and information retrieval. Conventional systems can only respond uniformly, ignoring the user's emotions, and it is difficult to respond flexibly according to the user's situation.

[1318] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1319] In this invention, the server includes means for analyzing the user's emotions from voice data and adjusting the content of proposals based on the emotions, means for analyzing the user's emotions when entering progress and generating a prompt alert if the urgency or importance is high, and means for adjusting the display content in consideration of the user's emotional state when generating search results. This enables flexible responses according to the user's emotional state, improving the quality and efficiency of sales activities.

[1320] Understood. Below are definitions of important words.

[1321] "Voice input" refers to the process by which a user inputs voice through a microphone.

[1322] "Convert to text" refers to the process of converting audio data into written information in real time.

[1323] "Generative AI" refers to artificial intelligence technology that learns from large amounts of data and generates new information and suggestions.

[1324] "Emotion analysis" refers to the process of analyzing a user's emotional state from input voice and text data.

[1325] "Emotion engine" refers to software or hardware used to recognize the emotional state of a user.

[1326] "Adjusting suggestions" refers to the process of changing the content of generated suggestions depending on the user's emotional state.

[1327] "Progress information" refers to data that indicates the progress or task status related to a project or sales activity.

[1328] "Progress management database" refers to a database for managing and storing progress information.

[1329] "Generating an alert" refers to the process of creating a notification to draw attention to an event of high urgency or importance.

[1330] "Search query" means a command or question entered by a user into a system in search of specific information.

[1331] "Search Results" refers to relevant information retrieved from a database based on a user's search query.

[1332] "Adjusting display content" refers to the process of changing the format and detail of the information displayed to take into account the user's emotional state.

[1333] MODE FOR CARRYING OUT THE INVENTION

[1334] The sales support system of the present invention converts a user's voice input into text in real time, and uses generative AI to summarize the information and make suggestions based on that data. It also incorporates an emotion engine that recognizes the user's emotions, and can adjust the content of suggestions based on the user's emotional state. A specific embodiment of this system is described below.

[1335] Hardware and software used

[1336] This system uses the following hardware and software:

[1337] Devices (e.g. smartphones, tablets, laptops)

[1338] Server (cloud server or on-premise server)

[1339] Voice recognition technology (e.g., Google Cloud Speech-to-Text)

[1340] Generative AI models (e.g., OpenAI's GPT-3)

[1341] Sentiment analysis engine (e.g., Microsoft Azure's Text Analytics API)

[1342] Specific Example of the System

[1343] Business negotiation support

[1344] 1. The user launches the app and uses voice input, for example, "What is the price of the new product?"

[1345] 2. The device converts the voice input into text in real time using Google Cloud Speech-to-Text technology and sends the text data to the server.

[1346] 3. The server analyzes the received text data using a generative AI model (GPT-3), extracts key points, and suggests the next action. For example, it generates a suggestion such as, "The price of this product is XX yen."

[1347] 4. The server uses an emotion analysis engine to analyze the user's emotions from the voice data, and if the user is nervous, adjusts the explanation or suggestions to be more friendly.

[1348] 5. The adjusted summary and suggestions are displayed to the user via the device.

[1349] Example prompt sentence:

[1350] "New product pricing information that can be used in business negotiations"

[1351] Progress management support

[1352] 1. The user inputs status information by voice or text, for example, "I need to schedule my next meeting with Client A."

[1353] 2. The device uses voice recognition technology to convert the voice into text and sends the data to the server.

[1354] 3. The server updates the progress management database and generates the next action or alert, for example, "A meeting has been scheduled with Client A."

[1355] 4. The server uses an emotion analysis engine to analyze the user's emotions when entering progress, and if it determines that the situation is urgent, it generates a prompt alert.

[1356] 5. The device notifies the user of the generated action or alert.

[1357] Example prompt sentence:

[1358] "Schedule my next meeting"

[1359] Information search support

[1360] 1. The user enters the information they are looking for by voice or text, for example, "What are the technical specifications for this product?"

[1361] 2. The device uses voice recognition technology to convert the voice into text and sends the data to the server.

[1362] 3. The server searches the database for relevant information and generates search results, such as "The technical specifications of this product are ____."

[1363] 4. When generating search results, the server analyzes the user's emotional state using an emotion analysis engine and adjusts the displayed content accordingly. For example, if the user is anxious, the server will provide information in a concise format.

[1364] 5. The device displays the adjusted search results to the user and provides audio feedback if necessary.

[1365] Example prompt sentence:

[1366] "What are the technical specifications of the product?"

[1367] This system aims to improve the quality and efficiency of sales activities by enabling flexible responses according to the user's emotional state.

[1368] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1369] Sales Support Processing Steps

[1370] Step 1:

[1371] The user launches the app and performs voice input.

[1372] Specifically, the user opens the app on their smartphone and asks, "What is the price of the new product?"

[1373] Input: User's voice data

[1374] Output: Audio input data

[1375] Step 2:

[1376] The device converts voice input into text in real time using speech recognition technology (e.g., Google Cloud Speech-to-Text).

[1377] Specifically, the device's microphone captures audio and converts it into text via a cloud service.

[1378] Input: Voice input data

[1379] Output: Text data

[1380] Step 3:

[1381] The terminal transmits the converted text data to the server.

[1382] Specifically, the device sends text data to the server over a secure HTTPS connection.

[1383] Input: Text data

[1384] Output: The text data sent

[1385] Step 4:

[1386] The server uses a generative AI model (e.g., GPT-3) to analyze the text data, extract key points, and suggest the next action.

[1387] Specifically, the server passes the text data to the analysis engine, and the generative AI model generates information about the "price of the new product."

[1388] Input: Text data sent

[1389] Output: Summary of proposal

[1390] Step 5:

[1391] The server uses a sentiment analysis engine (for example, Microsoft Azure's Text Analytics API) to analyze the user's emotions from the voice data and adjusts the suggestions based on those emotions.

[1392] Specifically, the server transfers voice data to an analysis engine, and if it detects that the user is in a tense state, it changes the suggestions to be more polite.

[1393] Input: Audio data, summarized proposal

[1394] Output: Sentiment-adjusted recommendations

[1395] Step 6:

[1396] The server sends the adjusted summary and suggestions to the terminal.

[1397] Specifically, the adjusted proposal is sent to the device over a secure connection.

[1398] Input: Emotion-adjusted recommendations

[1399] Output: Adjustment proposals submitted

[1400] Step 7:

[1401] The device displays the adjusted proposal to the user.

[1402] Specifically, the device will display on the screen, "The price of this product is XX yen. We are running a discount campaign for a limited time."

[1403] Input: Adjustment proposal submitted

[1404] Output: The suggestions displayed to the user

[1405] Progress management support processing steps

[1406] Step 1:

[1407] The user enters progress information by voice or text.

[1408] Specifically, the user might say or type in text, "I need to schedule my next meeting with Client A."

[1409] Input: Audio or text data

[1410] Output: Progress input data

[1411] Step 2:

[1412] The device converts the voice input into text using voice recognition technology and sends the data to the server.

[1413] Specifically, the terminal converts the progress input into text and securely transmits it to the server.

[1414] Input: Progress input data

[1415] Output: Progress data sent

[1416] Step 3:

[1417] The server updates the progress management database and generates the next action or alert.

[1418] As a specific operation, the server adds new progress information to the database and generates an action as "meeting set up."

[1419] Input: Submitted progress data

[1420] Output: Updated progress data, generated actions

[1421] Step 4:

[1422] The server uses an emotion analysis engine to analyze the user's emotions when entering progress information, and generates a prompt alert if the information is of high urgency or importance.

[1423] Specifically, the server analyzes the emotions expressed when entering progress, and if it determines that the situation is urgent, it generates a high-priority alert.

[1424] Input: Progress data, user emotion data

[1425] Output: The generated alert

[1426] Step 5:

[1427] The terminal notifies the user of the generated actions and alerts.

[1428] Specifically, the device will display a notification on the screen saying, "A meeting with Client A has been set up."

[1429] Input: Generated action, alert

[1430] Output: Information reported to the user

[1431] Information retrieval support processing steps

[1432] Step 1:

[1433] The user inputs the information they are looking for by voice or text.

[1434] Specifically, the user says, "Tell me the technical specifications of this product."

[1435] Input: Audio or text data

[1436] Output: Search query

[1437] Step 2:

[1438] The device converts the input into text using voice recognition technology and sends the data to the server.

[1439] Specifically, the device converts the voice input into text and sends it to the server.

[1440] Input: search query

[1441] Output: Submitted search query data

[1442] Step 3:

[1443] The server receives the query and retrieves the relevant information from a database.

[1444] Specifically, the server queries a product database to obtain the relevant technical specification information.

[1445] Input: Submitted search query data

[1446] Output: Related information

[1447] Step 4:

[1448] When generating search results, the server analyzes the user's emotional state using an emotion analysis engine and adjusts the display content.

[1449] Specifically, if the server determines that the user is in a hurry, it displays the technical specifications in a concise and easy-to-understand format.

[1450] Input: Related information, user emotion data

[1451] Output: Search results tailored based on sentiment

[1452] Step 5:

[1453] The device displays the adjusted search results to the user and provides audio feedback if necessary.

[1454] Specifically, the device will display on the screen, "The technical specifications of this product are XX. Please refer to this link for details," and the voice assistant will also read it out loud.

[1455] Input: Refined search results

[1456] Output: Search results displayed to the user

[1457] (Application example 2)

[1458] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1459] Sales activities in brick-and-mortar stores require immediate responses to customer questions and reactions, but conventional systems have difficulty taking customer emotions into account, often failing to provide appropriate information or suggestions. Furthermore, in situations where salespeople are at a loss as to how to respond, a system is needed that can generate and provide appropriate suggestions in real time. The present invention aims to solve these problems and provide a system that supports salespeople in effectively responding to customers.

[1460] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for using an emotion analysis engine that analyzes emotions from audio data and video data, means for generating summaries and proposals using generative AI, and means for the server to transmit summaries and proposals adjusted based on the emotion analysis. This makes it possible to provide information and proposals in real time that take customer emotions into consideration.

[1461] "User" refers to an individual or entity who uses the System to provide voice or text input.

[1462] A "terminal" is a device operated by a user that has the function of sending voice input and text data to a server.

[1463] "Server" refers to a central processing unit that receives text data sent from a device, processes the information using generative AI or an emotion analysis engine, and returns the results to the device.

[1464] "Generative AI" refers to an artificial intelligence engine that analyzes text data and automatically generates summaries of information and suggestions.

[1465] An "emotion analysis engine" is an engine that has the ability to analyze a user's emotions from audio and video data and adjust the content of suggestions based on that information.

[1466] "Voice input" refers to user voice data collected by the terminal, which is converted into text in real time.

[1467] "Text data" refers to information in which voice input is converted into text in real time.

[1468] A "summary" refers to information that is concisely summarized by generative AI that analyzes text data and extracts important points.

[1469] "Suggestion" refers to the next action or information the user should take, generated by generative AI based on text data.

[1470] "Progress management database" refers to a database for storing and managing user progress information.

[1471] An "alert" refers to urgent or important information that is notified to the user when the progress management database is updated.

[1472] "Search results" refers to information that the server searches for related information from the database and presents to the user.

[1473] The present invention is a system that supports sales activities in brick-and-mortar stores, and makes it possible to provide information and suggestions in real time based on the user's voice input and video data.

[1474] System Overview

[1475] The system includes the following main elements:

[1476] 1. User (salesperson): Operates the system and provides voice input.

[1477] 2. Terminal: A device operated by the user (e.g., smart glasses) that collects voice input and sends it to a server.

[1478] 3. Server: Analyzes data sent from the device and processes the information using generative AI and an emotion analysis engine.

[1479] Hardware and Software Details

[1480] Smart glasses: Devices with AR capabilities and microphones (e.g., Google Glass) that collect audio and visual data from the user.

[1481] Speech Recognition API: Uses the Google Cloud Speech-to-Text API to convert voice input to text in real time.

[1482] Generative AI model: Uses OpenAI GPT to analyze text data and generate information summaries and suggestions.

[1483] Sentiment analysis engine: Uses the Microsoft Azure Emotion API to analyze emotions from audio and video data.

[1484] Cloud server: Data processing is performed using Amazon Web Services (AWS).

[1485] Operation explanation

[1486] Receiving audio input

[1487] A user (salesperson) puts on the smart glasses during a sales meeting and starts a conversation with a customer. The microphone in the smart glasses collects the salesperson's voice input and temporarily stores the voice data. The stored voice data is converted into text data in real time by the Google Cloud Speech-to-Text API.

[1488] Sending and analyzing text data

[1489] The converted text data is sent to a cloud server (AWS). A generative AI model (OpenAI GPT) on the cloud server analyzes the text data, summarizes key points, and generates next actions and suggestions. In parallel, the audio and video data is analyzed by the Microsoft Azure Emotion API to determine the customer's emotions. Based on the results of the emotion analysis, the suggestions generated by the generative AI are adjusted appropriately.

[1490] Viewing Proposals

[1491] The final tailored summary and recommendations are then displayed on the smart glasses display, allowing the salesperson to respond appropriately to the customer.

[1492] Specific examples

[1493] For example, if a customer asks a salesperson, "How much does this product cost?", the smart glasses will collect the voice and convert it into text data in real time. This text data is sent to a server and analyzed by a generative AI model. The sentiment analysis engine will analyze the customer's emotions, and if the customer seems impatient, a short, concise, and unassuming answer will be generated. As a result, the salesperson can respond, "The price of this product is XX yen, but we are running a discount campaign for a limited time."

[1494] Prompt Sentence Examples

[1495] Convert user questions into text in real time, analyze customer sentiment, and generate optimal responses.

[1496] Example input:

[1497] Salesperson: "Look at this product, it has special features."

[1498] Customer: "What's the price?"

[1499] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1500] Step 1:

[1501] A user (salesperson) puts on smart glasses and starts a sales negotiation. As input, the microphone in the smart glasses collects voice data of the conversation with the customer. As output, the voice data is temporarily stored in the built-in memory. Specifically, the smart glasses continuously capture voice through the microphone.

[1502] Step 2:

[1503] The collected voice data is converted into text data in real time. As input, the voice data collected in step 1 is sent to the Google Cloud Speech-to-Text API. As output, the voice data is converted into text data. Specifically, the device (smart glasses) calls the Google Cloud Speech-to-Text API to convert the voice data into text.

[1504] Step 3:

[1505] The converted text data is sent to the server. As input, the text data converted in step 2 is sent to the cloud server (AWS). As output, the text data is saved on the server. In concrete terms, the device sends the text data to the cloud server via the Internet.

[1506] Step 4:

[1507] The server analyzes the text data and generates a summary and suggestions based on key points. The text data sent to the server in step 3 is received as input. The summary and suggestions are generated as output. Specifically, a generative AI model (OpenAI GPT) on the server analyzes the text data, extracts key elements, and generates suggestions.

[1508] Step 5:

[1509] In parallel, the audio and video data are sent to the emotion analysis engine on the server for analysis. As input, the audio data collected in step 1 and the video data captured by the smart glasses camera are sent to the server. As output, customer emotion data is generated. Specifically, the server uses the Microsoft Azure Emotion API to analyze emotions from the audio and video.

[1510] Step 6:

[1511] The generated proposals are adjusted based on the results of the sentiment analysis engine. The summaries and proposals generated in step 4 and the emotion data generated in step 5 are used as input. The output is a summary and proposal adjusted based on emotion. Specifically, the generative AI model on the server uses the emotion data to optimize the summaries and proposals.

[1512] Step 7:

[1513] The final summary and suggestions are sent to the terminal and displayed to the user. As input, the adjusted summary and suggestions generated in step 6 are sent from the cloud server to the terminal. As output, the suggestions are displayed on the display of the smart glasses. Specifically, the server sends the summary and suggestions to the smart glasses via the Internet, and the smart glasses display them.

[1514] Prompt Sentence Examples

[1515] Convert user questions into text in real time, analyze customer sentiment, and generate optimal responses.

[1516] Example input:

[1517] Salesperson: "Look at this product, it has special features."

[1518] Customer: "What's the price?"

[1519] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1520] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1521] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1522] [Fourth embodiment]

[1523] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1524] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1525] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1526] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1527] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1528] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1529] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1530] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1531] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1532] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1533] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1534] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1535] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1536] Understood. Below is the "Form for carrying out the invention".

[1537] In one embodiment of the invention, the system begins with a user providing instructions through voice input, which the device converts in real time into text and sends to the server. The server then uses generative AI to analyze the text data, generating summaries and suggestions, which are then sent back to the device for display to the user. When the user inputs progress, this is recorded in a progress management database, and the user is notified of next actions or alerts. Furthermore, when the user searches for specific information, the server searches for related data and provides the results to the user.

[1538] Business negotiation support

[1539] When a user launches the app, the app activates its voice input function. The device converts the user's speech into text in real time and sends the text data to a server. The server then uses generative AI to analyze the text data, summarize key points, and suggest next steps. These summaries and suggestions are displayed to the user on the device.

[1540] For example, suppose a user is introducing a new product during a business meeting and the client asks, "What is the price of this product?" In this case, the device recognizes the voice as a "price question" and sends the text data to the server. The server analyzes past data, generates appropriate price information, and returns it to the device. As a result, the user can respond to the client by saying, "The price of this product is XX yen."

[1541] Progress management support

[1542] In progress management, the user inputs the progress of each business negotiation or task by voice or text. The device recognizes this and sends the data to the server. The server updates the progress management database and generates the next action or alert. These actions and alerts are notified to the user via the device.

[1543] For example, when a user says, "Set up a meeting with Client A next week," the device converts the voice to text and sends it to the server. The server updates the progress management database and generates an alert to set up the next meeting. The device notifies the user when the meeting date approaches.

[1544] Information search support

[1545] With the information search function, users search for the information they need using voice or text. The device analyzes this input and sends the query to the server. The server then searches for relevant information in a database, generates results, and sends them to the device. The user can then view the search results through the device.

[1546] For example, if a user says, "I want to know the technical specifications of a new product," the device converts the instruction into text and sends it to the server. The server then searches the database for the new product's technical specifications and returns them to the device. The user can then view the specifications on the screen.

[1547] As a result, the system standardizes the quality of business negotiations, improves the efficiency of progress management, and enables quick and accurate information acquisition.

[1548] The processing flow will be explained below.

[1549] Understood. Below are the specific processing steps.

[1550] Business negotiation support

[1551] Step 1:

[1552] The user launches the app by tapping the app icon.

[1553] Step 2:

[1554] The terminal activates the voice input function and starts recording the user's voice in real time.

[1555] Step 3:

[1556] The device converts the recorded audio into text data in real time, using a speech recognition algorithm to output the audio as text.

[1557] Step 4:

[1558] The terminal transmits the text data to the server, and the text data is uploaded to the server via the network.

[1559] Step 5:

[1560] The server receives the text data and uses generative AI to summarize the information, including important points and next actions.

[1561] Step 6:

[1562] The server generates information summaries and suggestions, and creates optimal suggestions based on the results of analyzing the text data.

[1563] Step 7:

[1564] The server transmits the generated summary data and the proposal to the terminal, and transmits the data to the terminal via the network.

[1565] Step 8:

[1566] The device displays summary data and suggestions to the user, either as text on the screen or as audio feedback.

[1567] Progress management support

[1568] Step 1:

[1569] The user types a progress update into the app or gives a voice command. They tap the progress update button and enter the information.

[1570] Step 2:

[1571] The terminal performs voice recognition or text analysis on the input content to generate text data.

[1572] Step 3:

[1573] The device sends the text data to the server, where it is uploaded via the network.

[1574] Step 4:

[1575] The server receives the progress data and updates the progress management database, recording the new progress information.

[1576] Step 5:

[1577] The server analyzes the progress and generates the next action or alert: Create the necessary tasks or alerts.

[1578] Step 6:

[1579] The server sends the generated tasks and alerts to the terminal and returns the data via the network.

[1580] Step 7:

[1581] The device notifies the user of tasks and alerts, displaying them on the screen and providing audio feedback where appropriate.

[1582] Information search support

[1583] Step 1:

[1584] When a user wants to search for specific information, they input it by voice or text, then tap the search button.

[1585] Step 2:

[1586] The terminal performs voice recognition or text analysis on the input content to generate text data.

[1587] Step 3:

[1588] The device sends the text data to the server, where it is uploaded via the network.

[1589] Step 4:

[1590] The server receives the query and searches the database to extract information that matches the specified keywords.

[1591] Step 5:

[1592] The server generates search results and sends them to the device, organizes the results, and sends the data to the device.

[1593] Step 6:

[1594] The device displays the search results to the user, displaying them on the screen and providing audio feedback if necessary.

[1595] By following the above steps, this system provides functions such as support for business negotiations, progress management, and information search.

[1596] Example 1

[1597] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1598] Conventional sales negotiation support systems, progress management systems, and information search systems each function independently, resulting in low user convenience. Furthermore, there is a lack of systems that can efficiently convert voice input into text and then centrally manage the subsequent analysis, proposals, progress management, and information search. This creates challenges for users, making it difficult to standardize the quality of sales negotiations, efficiently manage progress, and quickly obtain information.

[1599] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1600] In this invention, the server includes means for analyzing information using a generative artificial intelligence model and generating summaries and proposals, means for updating the progress management database and generating next actions and alerts, and means for searching the database for related information, generating search results, and sending them to the terminal. This allows for centralized management of business negotiation support, progress management, and information search, enabling users to efficiently give instructions and obtain information.

[1601] "User" refers to any individual or corporation that uses this system.

[1602] "Voice input" refers to a means by which a user gives instructions to a system using voice.

[1603] "Terminal" refers to a hardware device that converts speech to text, transmits the data to a server, and displays the results.

[1604] "Server" refers to a computer system that uses generative artificial intelligence models to analyze text data and generate summaries, suggestions, progress management, and information search results.

[1605] A "generative artificial intelligence model" refers to an algorithm or program that analyzes input text data and generates summaries or suggestions based on it.

[1606] "Text data" refers to data resulting from converting voice input into text.

[1607] "Analyzing information" refers to the process of understanding input text data and generating summaries or suggestions based on that content.

[1608] A "summary" refers to a concise summary of the important points of text data.

[1609] "Suggestion" refers to showing the user the next action or countermeasure to be taken based on the analysis results.

[1610] A "progress management database" refers to a database for recording and managing the progress of business negotiations and tasks.

[1611] "Action" refers to the specific next steps or actions to be taken.

[1612] "Alert" refers to a warning or reminder that notifies the user of important information or progress.

[1613] "Information search" refers to a search performed by a user to obtain specific information.

[1614] "Database" refers to a data storage system in which related information is stored.

[1615] "Search Results" refers to a server-generated answer or collection of data in response to an information search query.

[1616] This invention is a system that unifies the management of business negotiation support, progress management, and information search. This system is composed of three main components: users, terminals, and servers.

[1617] Business negotiation support

[1618] When a user launches the app, the device activates the voice input function. When the user speaks, the device converts the speech into text in real time and sends the text data to the server. The server then analyzes the text data using a generative artificial intelligence model (e.g., OpenAI's GPT series), summarizes the key points, and suggests the next action to take. The suggested summary and action are displayed to the user on the device.

[1619] For example, if a user is introducing a new product and the client asks, "What is the price of this product?", the device converts the speech into text "Question about price" and sends it to the server. The server analyzes past data, generates appropriate price information, and returns it to the device. The user can then reply to the client, "The price of this product is XX yen."

[1620] Progress management

[1621] When the user inputs progress by voice or text, the device sends the data to the server, which updates the progress management database (e.g., PostgreSQL) and generates the next action or alert. The generated action or alert is then notified to the user via the device.

[1622] For example, if a user says, "Set up a meeting with Client A next week," the device converts the speech into text and sends it to the server. The server updates the progress management database and generates an alert to set up the next meeting. When the meeting date approaches, the device notifies the user.

[1623] Information Search

[1624] When a user searches for the information they need by voice or text, the device sends the query to the server, which then searches for relevant information in a database (e.g., Elasticsearch), generates results, and sends them to the device. The user can then view the search results through the device.

[1625] For example, if a user says, "I want to know the technical specifications of a new product," the device converts the instruction into text and sends it to the server. The server then searches the database for the new product's technical specifications and returns them to the device. The user can then view the specifications on the device's screen.

[1626] As a result, this system can standardize the quality of business negotiations, improve the efficiency of progress management, and enable quick and accurate information acquisition.

[1627] Prompt Sentence Examples

[1628] For sales support: "Please summarize the notes from your meeting with Client X and tell me what action to take next."

[1629] For progress management support: "How is project Y progressing? What's the next action?"

[1630] For information search assistance: "Please tell me the technical specifications for new product Z."

[1631] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1632] Understood. Below we will explain in detail the processing flow of the system program, divided into processing steps.

[1633] Business negotiation support

[1634] Step 1:

[1635] A user launches an app and activates the voice input function. The input is the command to launch the app, and the output is that voice input is enabled.

[1636] Step 2:

[1637] The user provides voice input, which the device converts into text in real time. The input is the user's voice data, and the output is text data. Specifically, the device uses voice recognition software (e.g., Google Speech-to-Text API).

[1638] Step 3:

[1639] The terminal sends the generated text data to the server. The input is the text data, and the output is the data sent to the server. Specifically, the terminal sends data to the server using an HTTPS request.

[1640] Step 4:

[1641] The server uses a generative AI model to analyze the transmitted text data. The input is text data, and the output is a summary and proposed data. Specifically, the server uses a generative AI model such as GPT-4 to perform the analysis and generation process.

[1642] Step 5:

[1643] The server sends the generated summary and proposal to the terminal. The input is the summary and proposal data, and the output is the data to be sent to the terminal. Specifically, the server sends the generated results in JSON format.

[1644] Step 6:

[1645] The terminal displays the received summary and suggestions to the user. The input is the data received from the server, and the output is the visualization to the user. Specifically, the terminal displays the data using a GUI.

[1646] Progress management

[1647] Step 1:

[1648] The user inputs their progress by voice or text. The input is the user's voice or text data, and the output is text data. Specifically, the device recognizes the voice and converts it into text.

[1649] Step 2:

[1650] The terminal sends the text data to the server. The input is the text data, and the output is the data sent to the server. Specific operations use an HTTPS request.

[1651] Step 3:

[1652] The server updates the progress management database. The input is text data, and the output is the updated database. Specifically, the server updates the database using an SQL query.

[1653] Step 4:

[1654] The server generates the next action or alert. The input is the updated database information, and the output is the action / alert data. Specifically, the server checks for pending tasks and generates the next action or alert.

[1655] Step 5:

[1656] The device notifies the user of the generated actions and alerts. The input is the action / alert data sent from the server, and the output is the notification to the user. Specifically, the device performs push notifications and in-app notifications.

[1657] Information Search

[1658] Step 1:

[1659] The user searches for the information they need by voice or text. The input is the user's query (voice or text), and the output is the query data. Specifically, the device performs voice recognition and converts it into text.

[1660] Step 2:

[1661] The device sends the query to the server. The input is the query data, and the output is the data sent to the server. Specific operations use an HTTPS request.

[1662] Step 3:

[1663] The server searches for relevant information from a database. The input is the query data, and the output is the search result data. Specifically, the server uses a search engine such as Elasticsearch.

[1664] Step 4:

[1665] The server generates search results and sends them to the terminal. The input is the search result data, and the output is the data to be sent to the terminal. Specifically, the results are sent in JSON format.

[1666] Step 5:

[1667] The terminal displays the search results to the user. The input is the data received from the server, and the output is the display to the user. Specifically, the terminal displays the data using a GUI.

[1668] (Application example 1)

[1669] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1670] Modern manufacturing factory operations require automation and real-time progress management. In particular, there is a lack of technology that allows factory operators to give work instructions to robots via voice and efficiently manage their progress, resulting in a decline in production efficiency and work accuracy. There is also a need for a system that can quickly search for and provide necessary information.

[1671] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1672] In this invention, the server includes: means for a user to give instructions through voice input and for the terminal to convert the voice into text in real time; means for sending text data to the server and for the server to generate a summary and proposal of the information using a generative AI model; means for the terminal to display the summary and proposal sent from the server to the user; means for the terminal to convert the voice into text and then send the text to the server; means for the server to generate appropriate actions and information in response to the instructions and send the results to the terminal; means for the user to input progress and tasks by voice or text and for the terminal to send them to the server; means for the server to update the progress management database and generate the next action or alert; and means for the terminal to notify the user of the generated action or alert. This enables factory operators to give instructions to robots by voice and quickly manage progress and obtain related information based on those instructions.

[1673] A "user" is a person who uses the system to give instructions and input progress.

[1674] "Voice input" is a method in which a user gives instructions or information to a terminal by voice.

[1675] A "terminal" is a device that converts voice input from a user into text and transmits the text data to a server.

[1676] "Real-time" is a time characteristic that means processing data and returning results almost instantaneously.

[1677] "Means for converting to text" refers to technology or devices for converting voice data into character data.

[1678] "Text data" is a data format in which voice is converted into text.

[1679] A "server" is a central computer system that receives text data and performs analysis and information generation.

[1680] A "generative AI model" is an algorithm that uses artificial intelligence technology to analyze received data and generate information summaries and suggestions.

[1681] "Means of summarizing and suggesting information" is the process of using a generative AI model to extract important information from the data received and present the user with the next action to take.

[1682] The "means for displaying the summary and suggestions" refers to a technique by which the terminal presents the summary and suggestions sent from the server to the user visually or audibly.

[1683] "Progress" is data that indicates the progress of a particular task or project.

[1684] The "progress management database" is a database system for storing and managing progress information entered by the user.

[1685] "Means for generating actions and alerts" refers to the process for generating next actions and alerts based on the progress management database.

[1686] "Means for notifying actions and alerts" refers to techniques for notifying users of generated actions and alerts.

[1687] "Means for searching information" refers to the technology that allows a user to send information searched for by voice or text to a server and return the results.

[1688] The "means for displaying search results" refers to a technique for visually or audibly presenting the search results sent from the server to the user.

[1689] The system required to implement this invention includes functions such as voice input, real-time text conversion, generative AI model, progress management, information search, etc. These will be explained in detail below.

[1690] Hardware used

[1691] The present invention uses the following hardware:

[1692] Microphone: Collects voice input from the user.

[1693] Device: A device that converts speech to text and communicates with the server. This can be a smartphone, tablet, or computer.

[1694] Server: Runs the generative AI model, analyzes data, and manages progress.

[1695] Software used

[1696] The following software is used in the present invention:

[1697] speech_recognition: A Python library for converting speech to text.

[1698] requests: A Python library for sending HTTP requests.

[1699] Generative AI models: Artificial intelligence techniques for analyzing text and generating summaries and suggestions.

[1700] TextToSpeech: A library for converting text to speech.

[1701] System Operation

[1702] 1. The user speaks a command into the device's microphone, for example, "Start the next shaft inspection."

[1703] 2. The device converts the voice input into text in real time and sends the text data to the server.

[1704] 3. The server uses a generative AI model to analyze the received text data and generate appropriate actions and information, such as a summary and suggestion such as "Initiate shaft inspection protocol."

[1705] 4. The generated actions and information are sent to the terminal and notified to the user by display or voice.

[1706] 5. When the user inputs their progress via voice or text, the device sends the data to the server, which updates the progress management database. The next action or alert is generated and notified to the user via the device.

[1707] Specific examples

[1708] Consider a factory operator who wants to perform a shaft inspection on a production line. The operator gives a voice command such as "Start the next shaft inspection." The device converts the voice to text and sends it to the server. The server uses a generative AI model to analyze the message and return it to the device: "Starting shaft inspection protocol." The operator proceeds with the work according to the protocol and reports the progress by voice. The server updates the progress management database and notifies the operator of the next task or alert.

[1709] Prompt Sentence Examples

[1710] Here are some examples of specific prompts:

[1711] "We will begin testing Shaft A next week."

[1712] "What are the technical specifications of the new product?"

[1713] This system allows factory operators to give voice instructions, which enable progress management and the rapid acquisition of related information.

[1714] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1715] Step 1:

[1716] The user inputs voice instructions through the microphone of the terminal, and the input voice data is sent to the terminal.

[1717] Step 2:

[1718] The device receives the voice data and converts it into text data in real time. Specifically, it uses the speech_recognition library to analyze the voice and output it as text data.

[1719] Step 3:

[1720] The text data generated by the terminal is sent to the server using an HTTP request.

[1721] Step 4:

[1722] The server inputs the received text data into a generative AI model for data analysis. The AI ​​model analyzes the instructions based on the prompt and generates a summary and proposal. For example, in response to the instruction "Start the next shaft inspection," it generates the summary "Start the shaft inspection protocol."

[1723] Step 5:

[1724] The server sends the generated summary and proposal to the terminal as an HTTP response.

[1725] Step 6:

[1726] The device receives the summary and suggestions from the server and displays them to the user, either by displaying the summary as text on the screen or by announcing it aloud using the TextToSpeech library.

[1727] Step 7:

[1728] The user inputs progress by voice or text, and the data is received by the terminal. For example, the user may report by voice that "the shaft inspection is complete."

[1729] Step 8:

[1730] The device converts the voice data into text in real time and sends the text data to the server, again using an HTTP request.

[1731] Step 9:

[1732] The server updates the progress management database based on the received text data. For example, it moves "Shaft Inspection" from the in-progress task list to the completed task list.

[1733] Step 10:

[1734] Based on the results of the updates to the progress management database, the server generates the next action or alert, based on pre-defined rules and conditions.

[1735] Step 11:

[1736] The server generates actions and alerts and sends them to the terminal as HTTP responses.

[1737] Step 12:

[1738] The device notifies the user of actions and alerts from the server, either by displaying them on the device screen or by voice using the TextToSpeech library.

[1739] Step 13:

[1740] When a user wants to search for information, the user inputs a query by voice or text, and the device receives the input. For example, the user may input a query such as "What are the technical specifications of a new product?"

[1741] Step 14:

[1742] The device converts the voice data into text and sends the text data to the server, again using an HTTP request.

[1743] Step 15:

[1744] The server searches the database based on the text data to retrieve relevant information, then analyzes the retrieved information using a generative AI model to generate search results in a format appropriate for the user.

[1745] Step 16:

[1746] The search results generated by the server are sent to the terminal as an HTTP response.

[1747] Step 17:

[1748] The device displays the search results from the server to the user, either as text on the screen or as audio using the TextToSpeech library.

[1749] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1750] Understood. Below is the "Form for carrying out the invention".

[1751] The sales support AI app of this invention is a system that converts user voice input into text in real time and uses that data to summarize information and make suggestions using generative AI. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and by adjusting the content of suggestions based on the user's emotional state, it provides more effective sales support.

[1752] Business negotiation support

[1753] During business negotiations, the user launches the app and uses voice input. The device converts the voice input into text in real time using voice recognition technology and sends the text data to the server. The server then uses generative AI to analyze the text data, summarizes key points, and suggests next steps. The server then uses an emotion engine to analyze the user's emotions from the voice data and adjusts the proposal content based on those emotions. The adjusted summary and proposal are then displayed to the user via the device.

[1754] For example, if a user is introducing a new product during a sales meeting and the customer asks, "How much does this product cost?", the device will recognize the voice, convert it into text data, and send it to the server. The server will analyze the text data and recognize the user's emotions through an emotion engine. If the user seems nervous, it will generate a more polite explanation and return it to the device. As a result, the user can respond, "The price of this product is XX yen, but we are running a discount campaign for a limited time."

[1755] Progress management support

[1756] To update progress, the user inputs progress information by voice or text. The device performs voice recognition or text analysis on this information and sends it to the server as text data. The server updates the progress management database and generates the next action or alert. The server also uses an emotion engine to analyze the user's emotions when entering progress information, and generates a faster alert if the information is urgent or important. The generated action or alert is then notified to the user via the device.

[1757] For example, if a user types, "I need to set up my next meeting with Client A," the device recognizes this, converts it into text data, and sends it to the server. The server updates the progress management database, analyzes the user's state of tension using an emotion engine, and then generates a high-priority alert and sends it to the device. As a result, the user receives a real-time notification that "A meeting with Client A has been set up."

[1758] Information search support

[1759] In information searches, users input the information they are looking for by voice or text. The device performs voice recognition or text analysis on the input and sends the text data to a server. The server receives the query and searches a database for relevant information. The emotion engine adjusts the content of the search results to take into account the user's emotions. For example, if the user is in a hurry, the device will display the search results in a concise summary format. The device then displays the results to the user and provides voice feedback as needed.

[1760] For example, if a user says, "Tell me the technical specifications of this product," the device converts this into text data and sends it to the server. The server searches the database for the relevant technical specifications, and if the emotion engine detects the user's impatience, it generates a concise and easy-to-understand technical specification and returns it to the device. As a result, the user can quickly and accurately tell the user, "The technical specifications of this product are ____, and for more details, please see this link."

[1761] In this way, by incorporating an emotion engine, it is possible to respond according to the user's emotional state, further improving the quality and efficiency of sales activities.

[1762] The processing flow will be explained below.

[1763] Understood. Below, I will explain in detail the processing steps of the sales support AI app that combines an emotion engine.

[1764] Business negotiation support

[1765] Step 1:

[1766] The user launches the app by tapping the app icon and begins voice input.

[1767] Step 2:

[1768] The device activates the voice input function and records the user's voice in real time, which is then sent to a voice recognition algorithm.

[1769] Step 3:

[1770] The device recognizes the recorded voice and converts it into text data, which is then temporarily saved.

[1771] Step 4:

[1772] The device sends text and voice data to the server in real time, where it is uploaded via a network connection.

[1773] Step 5:

[1774] The server receives the text data and uses generative AI to summarize the main points, while simultaneously analyzing the audio data with an emotion engine to recognize the user's emotional state.

[1775] Step 6:

[1776] Based on the analysis, the server generates suggestions that match the user's emotions, for example, if the user is nervous, a suggestion with more polite explanations will be generated.

[1777] Step 7:

[1778] The server then sends the generated summary data and sentiment-based suggestions to the device, where the data is transferred over the network.

[1779] Step 8:

[1780] The device displays summary data and suggestions to the user, providing on-screen text and, optionally, audio feedback.

[1781] Progress management support

[1782] Step 1:

[1783] The user enters progress information into the app, using voice or text input functionality to enter progress details.

[1784] Step 2:

[1785] The device converts voice into text in real time and temporarily stores the input text data.

[1786] Step 3:

[1787] The device sends the text data to the server, where it is uploaded over a network connection.

[1788] Step 4:

[1789] The server receives the progress data and updates the progress management database, and the new progress information is recorded in the database.

[1790] Step 5:

[1791] The server generates the next action or alert based on the progress data, using an emotion engine to analyze the user's emotional state and evaluate the urgency and importance of the action or alert.

[1792] Step 6:

[1793] The server sends the generated actions and alerts to the terminal, and the data is transferred over the network.

[1794] Step 7:

[1795] The device notifies the user of the generated actions and alerts by displaying them on the screen and / or providing audio feedback.

[1796] Information search support

[1797] Step 1:

[1798] The user enters the information they want to search for by voice or text, taps the search button, and enters search keywords.

[1799] Step 2:

[1800] The device converts voice input into text in real time and temporarily stores the text data.

[1801] Step 3:

[1802] The device sends text data to the server, where it is uploaded via the network.

[1803] Step 4:

[1804] The server searches the database based on the received query and extracts data that matches the specified keywords.

[1805] Step 5:

[1806] The server generates search results and analyzes the user's emotional state using an emotion engine, and adjusts the display format of the search results based on the user's emotion.

[1807] Step 6:

[1808] The server sends the adjusted search results to the device, and the data is transferred over the network.

[1809] Step 7:

[1810] The device displays the search results to the user, providing text on the screen and, optionally, audio feedback.

[1811] By following these steps, a sales support AI app incorporating an emotion engine will provide responses that correspond to the user's emotional state, improving the quality and efficiency of sales activities.

[1812] Example 2

[1813] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1814] In the processing of voice input in sales activities, the challenge is to generate appropriate information summaries and suggestions that take the user's emotions into account, thereby improving the efficiency of progress management and information retrieval. Conventional systems can only respond uniformly, ignoring the user's emotions, and it is difficult to respond flexibly according to the user's situation.

[1815] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1816] In this invention, the server includes means for analyzing the user's emotions from voice data and adjusting the content of proposals based on the emotions, means for analyzing the user's emotions when entering progress and generating a prompt alert if the urgency or importance is high, and means for adjusting the display content in consideration of the user's emotional state when generating search results. This enables flexible responses according to the user's emotional state, improving the quality and efficiency of sales activities.

[1817] Understood. Below are definitions of important words.

[1818] "Voice input" refers to the process by which a user inputs voice through a microphone.

[1819] "Convert to text" refers to the process of converting audio data into written information in real time.

[1820] "Generative AI" refers to artificial intelligence technology that learns from large amounts of data and generates new information and suggestions.

[1821] "Emotion analysis" refers to the process of analyzing a user's emotional state from input voice and text data.

[1822] "Emotion engine" refers to software or hardware used to recognize the emotional state of a user.

[1823] "Adjusting suggestions" refers to the process of changing the content of generated suggestions depending on the user's emotional state.

[1824] "Progress information" refers to data that indicates the progress or task status related to a project or sales activity.

[1825] "Progress management database" refers to a database for managing and storing progress information.

[1826] "Generating an alert" refers to the process of creating a notification to draw attention to an event of high urgency or importance.

[1827] "Search query" means a command or question entered by a user into a system in search of specific information.

[1828] "Search Results" refers to relevant information retrieved from a database based on a user's search query.

[1829] "Adjusting display content" refers to the process of changing the format and detail of the information displayed to take into account the user's emotional state.

[1830] MODE FOR CARRYING OUT THE INVENTION

[1831] The sales support system of the present invention converts a user's voice input into text in real time, and uses generative AI to summarize the information and make suggestions based on that data. It also incorporates an emotion engine that recognizes the user's emotions, and can adjust the content of suggestions based on the user's emotional state. A specific embodiment of this system is described below.

[1832] Hardware and software used

[1833] This system uses the following hardware and software:

[1834] Devices (e.g. smartphones, tablets, laptops)

[1835] Server (cloud server or on-premise server)

[1836] Voice recognition technology (e.g., Google Cloud Speech-to-Text)

[1837] Generative AI models (e.g., OpenAI's GPT-3)

[1838] Sentiment analysis engine (e.g., Microsoft Azure's Text Analytics API)

[1839] Specific Example of the System

[1840] Business negotiation support

[1841] 1. The user launches the app and uses voice input, for example, "What is the price of the new product?"

[1842] 2. The device converts the voice input into text in real time using Google Cloud Speech-to-Text technology and sends the text data to the server.

[1843] 3. The server analyzes the received text data using a generative AI model (GPT-3), extracts key points, and suggests the next action. For example, it generates a suggestion such as, "The price of this product is XX yen."

[1844] 4. The server uses an emotion analysis engine to analyze the user's emotions from the voice data, and if the user is nervous, adjusts the explanation or suggestions to be more friendly.

[1845] 5. The adjusted summary and suggestions are displayed to the user via the device.

[1846] Example prompt sentence:

[1847] "New product pricing information that can be used in business negotiations"

[1848] Progress management support

[1849] 1. The user inputs status information by voice or text, for example, "I need to schedule my next meeting with Client A."

[1850] 2. The device uses voice recognition technology to convert the voice into text and sends the data to the server.

[1851] 3. The server updates the progress management database and generates the next action or alert, for example, "A meeting has been scheduled with Client A."

[1852] 4. The server uses an emotion analysis engine to analyze the user's emotions when entering progress, and if it determines that the situation is urgent, it generates a prompt alert.

[1853] 5. The device notifies the user of the generated action or alert.

[1854] Example prompt sentence:

[1855] "Schedule my next meeting"

[1856] Information search support

[1857] 1. The user enters the information they are looking for by voice or text, for example, "What are the technical specifications for this product?"

[1858] 2. The device uses voice recognition technology to convert the voice into text and sends the data to the server.

[1859] 3. The server searches the database for relevant information and generates search results, such as "The technical specifications of this product are ____."

[1860] 4. When generating search results, the server analyzes the user's emotional state using an emotion analysis engine and adjusts the displayed content accordingly. For example, if the user is anxious, the server will provide information in a concise format.

[1861] 5. The device displays the adjusted search results to the user and provides audio feedback if necessary.

[1862] Example prompt sentence:

[1863] "What are the technical specifications of the product?"

[1864] This system aims to improve the quality and efficiency of sales activities by enabling flexible responses according to the user's emotional state.

[1865] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1866] Sales Support Processing Steps

[1867] Step 1:

[1868] The user launches the app and performs voice input.

[1869] Specifically, the user opens the app on their smartphone and asks, "What is the price of the new product?"

[1870] Input: User's voice data

[1871] Output: Audio input data

[1872] Step 2:

[1873] The device converts voice input into text in real time using speech recognition technology (e.g., Google Cloud Speech-to-Text).

[1874] Specifically, the device's microphone captures audio and converts it into text via a cloud service.

[1875] Input: Voice input data

[1876] Output: Text data

[1877] Step 3:

[1878] The terminal transmits the converted text data to the server.

[1879] Specifically, the device sends text data to the server over a secure HTTPS connection.

[1880] Input: Text data

[1881] Output: The text data sent

[1882] Step 4:

[1883] The server uses a generative AI model (e.g., GPT-3) to analyze the text data, extract key points, and suggest the next action.

[1884] Specifically, the server passes the text data to the analysis engine, and the generative AI model generates information about the "price of the new product."

[1885] Input: Text data sent

[1886] Output: Summary of proposal

[1887] Step 5:

[1888] The server uses a sentiment analysis engine (for example, Microsoft Azure's Text Analytics API) to analyze the user's emotions from the voice data and adjusts the suggestions based on those emotions.

[1889] Specifically, the server transfers voice data to an analysis engine, and if it detects that the user is in a tense state, it changes the suggestions to be more polite.

[1890] Input: Audio data, summarized proposal

[1891] Output: Sentiment-adjusted recommendations

[1892] Step 6:

[1893] The server sends the adjusted summary and suggestions to the terminal.

[1894] Specifically, the adjusted proposal is sent to the device over a secure connection.

[1895] Input: Emotion-adjusted recommendations

[1896] Output: Adjustment proposals submitted

[1897] Step 7:

[1898] The device displays the adjusted proposal to the user.

[1899] Specifically, the device will display on the screen, "The price of this product is XX yen. We are running a discount campaign for a limited time."

[1900] Input: Adjustment proposal submitted

[1901] Output: The suggestions displayed to the user

[1902] Progress management support processing steps

[1903] Step 1:

[1904] The user enters progress information by voice or text.

[1905] Specifically, the user might say or type in text, "I need to schedule my next meeting with Client A."

[1906] Input: Audio or text data

[1907] Output: Progress input data

[1908] Step 2:

[1909] The device converts the voice input into text using voice recognition technology and sends the data to the server.

[1910] Specifically, the terminal converts the progress input into text and securely transmits it to the server.

[1911] Input: Progress input data

[1912] Output: Progress data sent

[1913] Step 3:

[1914] The server updates the progress management database and generates the next action or alert.

[1915] As a specific operation, the server adds new progress information to the database and generates an action as "meeting set up."

[1916] Input: Submitted progress data

[1917] Output: Updated progress data, generated actions

[1918] Step 4:

[1919] The server uses an emotion analysis engine to analyze the user's emotions when entering progress information, and generates a prompt alert if the information is of high urgency or importance.

[1920] Specifically, the server analyzes the emotions expressed when entering progress, and if it determines that the situation is urgent, it generates a high-priority alert.

[1921] Input: Progress data, user emotion data

[1922] Output: The generated alert

[1923] Step 5:

[1924] The terminal notifies the user of the generated actions and alerts.

[1925] Specifically, the device will display a notification on the screen saying, "A meeting with Client A has been set up."

[1926] Input: Generated action, alert

[1927] Output: Information reported to the user

[1928] Information retrieval support processing steps

[1929] Step 1:

[1930] The user inputs the information they are looking for by voice or text.

[1931] Specifically, the user says, "Tell me the technical specifications of this product."

[1932] Input: Audio or text data

[1933] Output: Search query

[1934] Step 2:

[1935] The device converts the input into text using voice recognition technology and sends the data to the server.

[1936] Specifically, the device converts the voice input into text and sends it to the server.

[1937] Input: search query

[1938] Output: Submitted search query data

[1939] Step 3:

[1940] The server receives the query and retrieves the relevant information from a database.

[1941] Specifically, the server queries a product database to obtain the relevant technical specification information.

[1942] Input: Submitted search query data

[1943] Output: Related information

[1944] Step 4:

[1945] When generating search results, the server analyzes the user's emotional state using an emotion analysis engine and adjusts the display content.

[1946] Specifically, if the server determines that the user is in a hurry, it displays the technical specifications in a concise and easy-to-understand format.

[1947] Input: Related information, user emotion data

[1948] Output: Search results tailored based on sentiment

[1949] Step 5:

[1950] The device displays the adjusted search results to the user and provides audio feedback if necessary.

[1951] Specifically, the device will display on the screen, "The technical specifications of this product are XX. Please refer to this link for details," and the voice assistant will also read it out loud.

[1952] Input: Refined search results

[1953] Output: Search results displayed to the user

[1954] (Application example 2)

[1955] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1956] Sales activities in brick-and-mortar stores require immediate responses to customer questions and reactions, but conventional systems have difficulty taking customer emotions into account, often failing to provide appropriate information or suggestions. Furthermore, in situations where salespeople are at a loss as to how to respond, a system is needed that can generate and provide appropriate suggestions in real time. The present invention aims to solve these problems and provide a system that supports salespeople in effectively responding to customers.

[1957] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for using an emotion analysis engine that analyzes emotions from audio data and video data, means for generating summaries and proposals using generative AI, and means for the server to transmit summaries and proposals adjusted based on the emotion analysis. This makes it possible to provide information and proposals in real time that take customer emotions into consideration.

[1958] "User" refers to an individual or entity who uses the System to provide voice or text input.

[1959] A "terminal" is a device operated by a user that has the function of sending voice input and text data to a server.

[1960] "Server" refers to a central processing unit that receives text data sent from a device, processes the information using generative AI or an emotion analysis engine, and returns the results to the device.

[1961] "Generative AI" refers to an artificial intelligence engine that analyzes text data and automatically generates summaries of information and suggestions.

[1962] An "emotion analysis engine" is an engine that has the ability to analyze a user's emotions from audio and video data and adjust the content of suggestions based on that information.

[1963] "Voice input" refers to user voice data collected by the terminal, which is converted into text in real time.

[1964] "Text data" refers to information in which voice input is converted into text in real time.

[1965] A "summary" refers to information that is concisely summarized by generative AI that analyzes text data and extracts important points.

[1966] "Suggestion" refers to the next action or information the user should take, generated by generative AI based on text data.

[1967] "Progress management database" refers to a database for storing and managing user progress information.

[1968] An "alert" refers to urgent or important information that is notified to the user when the progress management database is updated.

[1969] "Search results" refers to information that the server searches for related information from the database and presents to the user.

[1970] The present invention is a system that supports sales activities in brick-and-mortar stores, and makes it possible to provide information and suggestions in real time based on the user's voice input and video data.

[1971] System Overview

[1972] The system includes the following main elements:

[1973] 1. User (salesperson): Operates the system and provides voice input.

[1974] 2. Terminal: A device operated by the user (e.g., smart glasses) that collects voice input and sends it to a server.

[1975] 3. Server: Analyzes data sent from the device and processes the information using generative AI and an emotion analysis engine.

[1976] Hardware and Software Details

[1977] Smart glasses: Devices with AR capabilities and microphones (e.g., Google Glass) that collect audio and visual data from the user.

[1978] Speech Recognition API: Uses the Google Cloud Speech-to-Text API to convert voice input to text in real time.

[1979] Generative AI model: Uses OpenAI GPT to analyze text data and generate information summaries and suggestions.

[1980] Sentiment analysis engine: Uses the Microsoft Azure Emotion API to analyze emotions from audio and video data.

[1981] Cloud server: Data processing is performed using Amazon Web Services (AWS).

[1982] Operation explanation

[1983] Receiving audio input

[1984] A user (salesperson) puts on the smart glasses during a sales meeting and starts a conversation with a customer. The microphone in the smart glasses collects the salesperson's voice input and temporarily stores the voice data. The stored voice data is converted into text data in real time by the Google Cloud Speech-to-Text API.

[1985] Sending and analyzing text data

[1986] The converted text data is sent to a cloud server (AWS). A generative AI model (OpenAI GPT) on the cloud server analyzes the text data, summarizes key points, and generates next actions and suggestions. In parallel, the audio and video data is analyzed by the Microsoft Azure Emotion API to determine the customer's emotions. Based on the results of the emotion analysis, the suggestions generated by the generative AI are adjusted appropriately.

[1987] Viewing Proposals

[1988] The final tailored summary and recommendations are then displayed on the smart glasses display, allowing the salesperson to respond appropriately to the customer.

[1989] Specific examples

[1990] For example, if a customer asks a salesperson, "How much does this product cost?", the smart glasses will collect the voice and convert it into text data in real time. This text data is sent to a server and analyzed by a generative AI model. The sentiment analysis engine will analyze the customer's emotions, and if the customer seems impatient, a short, concise, and unassuming answer will be generated. As a result, the salesperson can respond, "The price of this product is XX yen, but we are running a discount campaign for a limited time."

[1991] Prompt Sentence Examples

[1992] Convert user questions into text in real time, analyze customer sentiment, and generate optimal responses.

[1993] Example input:

[1994] Salesperson: "Look at this product, it has special features."

[1995] Customer: "What's the price?"

[1996] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1997] Step 1:

[1998] A user (salesperson) puts on smart glasses and starts a sales negotiation. As input, the microphone in the smart glasses collects voice data of the conversation with the customer. As output, the voice data is temporarily stored in the built-in memory. Specifically, the smart glasses continuously capture voice through the microphone.

[1999] Step 2:

[2000] The collected voice data is converted into text data in real time. As input, the voice data collected in step 1 is sent to the Google Cloud Speech-to-Text API. As output, the voice data is converted into text data. Specifically, the device (smart glasses) calls the Google Cloud Speech-to-Text API to convert the voice data into text.

[2001] Step 3:

[2002] The converted text data is sent to the server. As input, the text data converted in step 2 is sent to the cloud server (AWS). As output, the text data is saved on the server. In concrete terms, the device sends the text data to the cloud server via the Internet.

[2003] Step 4:

[2004] The server analyzes the text data and generates a summary and suggestions based on key points. The text data sent to the server in step 3 is received as input. The summary and suggestions are generated as output. Specifically, a generative AI model (OpenAI GPT) on the server analyzes the text data, extracts key elements, and generates suggestions.

[2005] Step 5:

[2006] In parallel, the audio and video data are sent to the emotion analysis engine on the server for analysis. As input, the audio data collected in step 1 and the video data captured by the smart glasses camera are sent to the server. As output, customer emotion data is generated. Specifically, the server uses the Microsoft Azure Emotion API to analyze emotions from the audio and video.

[2007] Step 6:

[2008] The generated proposals are adjusted based on the results of the sentiment analysis engine. The summaries and proposals generated in step 4 and the emotion data generated in step 5 are used as input. The output is a summary and proposal adjusted based on emotion. Specifically, the generative AI model on the server uses the emotion data to optimize the summaries and proposals.

[2009] Step 7:

[2010] The final summary and suggestions are sent to the terminal and displayed to the user. As input, the adjusted summary and suggestions generated in step 6 are sent from the cloud server to the terminal. As output, the suggestions are displayed on the display of the smart glasses. Specifically, the server sends the summary and suggestions to the smart glasses via the Internet, and the smart glasses display them.

[2011] Prompt Sentence Examples

[2012] Convert user questions into text in real time, analyze customer sentiment, and generate optimal responses.

[2013] Example input:

[2014] Salesperson: "Look at this product, it has special features."

[2015] Customer: "What's the price?"

[2016] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2017] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2018] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2019] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2020] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2021] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2022] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2023] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2024] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2025] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2026] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2027] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2028] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2029] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2030] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2031] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2032] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2033] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2034] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2035] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2036] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2037] The following is further disclosed regarding the above embodiment.

[2038] Understood. Below is a proposed draft of the claims for the patent application for the sales support AI app.

[2039] (Claim 1)

[2040] a means for a user to provide instructions through voice input, the device converting the speech into text in real time;

[2041] a means for transmitting the text data to a server, and the server using generative AI to generate summaries and suggestions for the information;

[2042] means for the terminal to display to the user the summaries and suggestions sent from the server;

[2043] A system including:

[2044] (Claim 2)

[2045] A means for the user to input progress by voice or text and for the device to transmit it to the server;

[2046] A means for the server to update the progress management database and generate next actions and alerts;

[2047] means for the terminal to notify the user of the generated action or alert;

[2048] 10. The system of claim 1, comprising:

[2049] (Claim 3)

[2050] A means for a user to search for specific information by voice or text, and for the terminal to transmit the information to a server;

[2051] A means for the server to search for relevant information from the database, generate search results, and transmit them to the terminal;

[2052] means for the terminal to display search results to the user;

[2053] 10. The system of claim 1, comprising:

[2054] "Example 1"

[2055] (Claim 1)

[2056] a means for a user to provide instructions through voice input, the device converting the speech into text in real time;

[2057] means for transmitting the text data to a server, which uses a generative artificial intelligence model to analyze the information and generate summaries and suggestions;

[2058] means for the terminal to display to the user the summaries and suggestions sent from the server;

[2059] a means for the user to input progress by voice or text and for the device to transmit the data to a server;

[2060] A means for the server to update the progress management database and generate next actions and alerts;

[2061] means for the terminal to notify the user of the generated action or alert;

[2062] A means for a user to search for specific information by voice or text, and for the terminal to transmit the information to a server;

[2063] A means for the server to search for relevant information from the database, generate search results, and transmit them to the terminal;

[2064] means for the terminal to display search results to the user;

[2065] A system including:

[2066] (Claim 2)

[2067] A means for the server to analyze past data and generate appropriate information based on voice instructions;

[2068] A means for the server to generate actions and alerts from the progress management database and notify the terminal;

[2069] 10. The system of claim 1, comprising:

[2070] (Claim 3)

[2071] means for searching the database for relevant information after the server receives a specific information search instruction, generating search results and transmitting the results to the terminal;

[2072] means for the terminal to display search results to the user;

[2073] 10. The system of claim 1, comprising:

[2074] "Application Example 1"

[2075] (Claim 1)

[2076] a means for a user to provide instructions through voice input, the device converting the speech into text in real time;

[2077] means for transmitting the text data to a server, which uses a generative AI model to generate summaries and suggestions for the information;

[2078] means for the terminal to display to the user the summaries and suggestions sent from the server;

[2079] means for converting speech to text by the device and then transmitting the text to a server;

[2080] A means for the server to generate appropriate actions and information in response to the instruction and transmit the results to the terminal;

[2081] A means for the user to input progress and tasks by voice or text, and the device sends it to the server;

[2082] A means for the server to update the progress management database and generate next actions and alerts;

[2083] means for the terminal to notify the user of the generated action or alert;

[2084] A system including:

[2085] (Claim 2)

[2086] A means for a user to search for specific information by voice or text, and for the terminal to transmit the information to a server;

[2087] A means for the server to search for relevant information from the database, generate search results, and transmit them to the terminal;

[2088] means for the terminal to display search results to the user;

[2089] 10. The system of claim 1, comprising:

[2090] (Claim 3)

[2091] means for the terminal to audibly notify the user of the generated summary and suggestions;

[2092] 10. The system of claim 1, comprising:

[2093] "Example 2: Combining Emotion Engines"

[2094] New Claims

[2095] (Claim 1)

[2096] a means for a user to provide instructions through voice input, the device converting the speech into text in real time;

[2097] a means for transmitting the text data to a server, and the server using generative AI to generate summaries and suggestions for the information;

[2098] A means for the server to analyze the user's emotions from the voice data and adjust the content of the suggestions based on the emotions;

[2099] means for the terminal to display to the user the summary and tailored proposals sent from the server;

[2100] A system including:

[2101] (Claim 2)

[2102] A means for the user to input progress by voice or text and for the device to transmit it to the server;

[2103] A means for the server to update the progress management database and generate next actions and alerts;

[2104] The server analyzes the user's emotions when entering progress, and generates a prompt alert if the situation is urgent or important.

[2105] means for the terminal to notify the user of the generated action or alert;

[2106] 10. The system of claim 1, comprising:

[2107] (Claim 3)

[2108] A means for a user to search for specific information by voice or text, and for the terminal to transmit the information to a server;

[2109] A means for the server to search for relevant information from the database, generate search results, and transmit them to the terminal;

[2110] a means for adjusting the display content by the server when generating search results, taking into account the emotional state of the user;

[2111] means for the device to display tailored search results to the user;

[2112] 10. The system of claim 1, comprising:

[2113] "Application example 2 when combining emotion engines"

[2114] (Claim 1)

[2115] a means for a user to provide instructions through voice input, the device converting the speech into text in real time;

[2116] a means for transmitting the text data to a server, and the server using generative AI to generate summaries and suggestions for the information;

[2117] A server uses an emotion analysis engine to analyze emotions from audio data and video data;

[2118] means for the terminal to display to the user summaries and suggestions tailored based on the sentiment analysis sent from the server;

[2119] A system including:

[2120] (Claim 2)

[2121] A means for the user to input progress by voice or text and for the device to transmit it to the server;

[2122] A means for the server to update the progress management database and generate next actions and alerts;

[2123] A means for the server to use an emotion analysis engine that analyzes emotions from audio and video data when inputting progress and determines the urgency and importance of the progress;

[2124] means for the terminal to notify the user of the generated action or alert;

[2125] 10. The system of claim 1, comprising:

[2126] (Claim 3)

[2127] A means for a user to search for specific information by voice or text, and for the terminal to transmit the information to a server;

[2128] A means for the server to search for relevant information from the database, generate search results, and transmit them to the terminal;

[2129] A server uses an emotion analysis engine to analyze user emotions from audio data and video data and adjust search results;

[2130] means for the device to display tailored search results to the user;

[2131] 10. The system of claim 1, comprising: [Explanation of symbols]

[2132] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for a user to provide instructions through voice input, the device converting the speech into text in real time; a means for transmitting the text data to a server, and the server using generative AI to generate summaries and suggestions for the information; means for the terminal to display to the user the summaries and suggestions sent from the server; A system including:

2. A means for the user to input progress by voice or text and for the device to transmit it to the server; A means for the server to update the progress management database and generate next actions and alerts; means for the terminal to notify the user of the generated action or alert; The system of claim 1 , comprising:

3. A means for a user to search for specific information by voice or text, and for the terminal to transmit the information to a server; A means for the server to search for relevant information from the database, generate search results, and transmit them to the terminal; means for the terminal to display search results to the user; The system of claim 1 , comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A