system
A resident generative AI system addresses inefficiencies in information processing and task management by converting audio to text, analyzing user input, and generating real-time responses, enhancing work efficiency and reducing stress.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-08
AI Technical Summary
Users face inefficiencies and increased stress due to the time and effort required for information search, text conversion, task management, email correspondence, and function input, exacerbated by generational and preference gaps, leading to delays and confusion in work tasks.
A system utilizing a resident generative AI model that collects audio data, converts it to text, analyzes user input, searches for word meanings, generates preliminary replies, suggests functions, and manages tasks, all in real-time to streamline these processes.
Significantly reduces the time and effort needed for information retrieval, data entry, and task management, enhancing work efficiency and reducing stress by providing real-time assistance.
Smart Images

Figure 2026060627000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern information society, users are faced with a large amount of information in their work and daily life. Along with this, they spend a lot of time and effort on tasks such as information search, text conversion, task management, email correspondence, function input, etc. In such a situation, the work efficiency of users decreases and stress increases. Also, there is a gap in information understanding due to differences between different generations and differences in hobbies and preferences, which further causes delays and confusion in work. Furthermore, accuracy and efficiency in daily work such as recording meeting content and task reminders are also required.
Means for Solving the Problems
[0005] To solve this problem, the present invention provides a system that utilizes a resident generative AI model to semi-automatically or automatically respond to various needs in the user's work and daily life. The system of the present invention includes the following means:
[0006] 1. A means of collecting audio data through the microphone of a device used by the user and transmitting the collected audio data to a server.
[0007] 2. A means by which a server analyzes audio data, converts it to text, and displays the converted text on a terminal in real time.
[0008] 3. A means of analyzing text entered or selected by the user on the device to detect unknown words or phrases.
[0009] 4. A means by which the server searches for the meaning and related information of detected words and phrases, and displays the search results on the device as a pop-up notification.
[0010] 5. A means by which the server analyzes the content of a received email, generates a preliminary reply, and displays the generated preliminary reply as a draft on the terminal.
[0011] 6. A method for analyzing the content entered by the user into a spreadsheet, suggesting appropriate functions or scripts, and generating and automatically inserting the suggested functions or scripts into cells.
[0012] 7. A means for users to input tasks into a scheduler or task management app, and a means for transmitting the input task data to a server in real time.
[0013] 8. A method by which the server analyzes task data, identifies and reminds users of uncompleted tasks with deadlines, and displays reminder notifications to users at an appropriate time.
[0014] This allows users to significantly reduce the time and effort required for searching for and entering information, managing tasks, and other related activities in their work and daily lives, enabling them to work more efficiently and effectively.
[0015] A "terminal" is a device that a user directly operates to input information or give instructions, and includes computers, smartphones, tablets, and other similar devices.
[0016] A "server" is a central processing unit that receives data sent from terminals and performs processing such as analysis, transformation, and retrieval.
[0017] "Audio data" refers to data that digitally represents audio from meetings, conversations, and other similar events.
[0018] "Text" refers to written material that is obtained by analyzing audio data and converting it into written form.
[0019] "Analysis" is the act of investigating and breaking down received data to extract information and process it according to a specific purpose.
[0020] "Conversion" refers to the process of changing data in one format to another. In this invention, it mainly refers to converting audio data into text.
[0021] A "pop-up notification" is a message that temporarily appears on the screen the user is using, serving as a means of instantly informing them of important information or suggestions.
[0022] "Initial reply" refers to the initial response to a received email, which is automatically generated by the server.
[0023] A "function" refers to a mathematical formula used to calculate or process specific numbers or data, and is particularly common in spreadsheets and Excel.
[0024] A "script" is a set of instructions written to automate a specific task.
[0025] A "task" refers to a specific operation or work item that a user should perform.
[0026] A "reminder" is an act of notifying a user of a specific task or other important matter to prompt re - recognition.
Brief Description of Drawings
[0027] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset - type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0028] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0029] First, let's explain the terminology used in the following explanation.
[0030] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).
[0031] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0032] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0033] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0034] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0035] [First Embodiment]
[0036] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0037] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0038] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0039] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0040] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0041] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0042] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0043] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0044] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0045] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0046] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0047] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0048] This invention relates to a resident generation AI system for streamlining information retrieval, text conversion, task management, email correspondence, function input, and other tasks in users' work and daily lives. This invention comprises the following components and their operation.
[0049] System Configuration
[0050] Real-time transcription of meeting content
[0051] Collection and transmission of audio data
[0052] Terminal: Collects audio in real time via microphone during meetings and conversations.
[0053] Terminal: Digitizes the collected audio data and sends it to the server.
[0054] Audio data analysis and text conversion
[0055] Server: Inputs the received audio data into the speech recognition model and converts it into text.
[0056] Server: Temporarily stores the text-based data and provides the user with appropriately filtered content.
[0057] Display text
[0058] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[0059] User: Review the displayed text and edit or save it as needed.
[0060] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[0061] Notification of the meaning of unfamiliar words
[0062] Text input and analysis
[0063] User: Enter text into Excel sheets or documents.
[0064] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[0065] Word meaning search
[0066] Server: Searches for the meaning of a word or phrase selected by the user from related databases.
[0067] Server: Analyzes relevant information and extracts appropriate content.
[0068] Meaning notification
[0069] Device: Display search results as a pop-up notification.
[0070] User: Check the displayed explanation and quickly obtain the necessary information.
[0071] Specific example: If a user doesn't understand the meaning of a particular function in Excel, selecting that function will display its meaning in a pop-up on the device. This allows the user to understand the meaning without interrupting their work.
[0072] Suggestion for an automated reply
[0073] Receiving and analyzing emails
[0074] Terminal: Notifies the user that a new email has been received and displays its contents.
[0075] Server: Analyzes the content of the email and generates an appropriate initial reply.
[0076] Reply generation and display
[0077] Server: Generates an initial reply and applies it to the template.
[0078] Terminal: Displays the generated initial reply as a draft.
[0079] User: Review the draft, make any necessary corrections, and then send a reply.
[0080] Specific example: When a user receives an email, the server analyzes its contents and generates an initial reply such as, "Could you please schedule a meeting on the following dates?" The user can then review and send the email, allowing for a quick response.
[0081] Function suggestions for Excel and spreadsheets
[0082] Text input and analysis
[0083] User: Enter data into Excel or a spreadsheet.
[0084] Terminal: Analyzes the input data and suggests appropriate functions or scripts.
[0085] Function proposal generation and input assistance
[0086] Server: Based on the input, it selects the most appropriate function or script and generates it as a suggestion.
[0087] Terminal: Automatically inserts the suggested function into the cell.
[0088] User: Review the proposed function and set the required data range and conditions.
[0089] Specific example: When a user aggregates sales data, the "SUM function" is suggested, and the terminal automatically inserts the function. The user can perform the aggregation simply by specifying the data range.
[0090] Task organization and reminder function
[0091] Task data collection and transmission
[0092] User: Enter tasks into a task management app or scheduler.
[0093] Terminal: Sends entered task data to the server in real time.
[0094] Task analysis and reminders
[0095] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[0096] Server: Creates reminder notifications for tasks with approaching deadlines.
[0097] Display reminder notifications
[0098] Device: Displays reminder notifications to the user at the appropriate time.
[0099] User: Check the reminder notification and take action on the task.
[0100] Specific example: When a user uses a task management app and the deadline approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user can then see the reminder and efficiently work on the task.
[0101] The system of the present invention allows users to efficiently carry out various tasks in their work and daily lives, significantly reducing the time and effort required for information retrieval, data entry, task management, and other related activities.
[0102] The following describes the processing flow.
[0103] Real-time transcription of meeting content
[0104] Step 1:
[0105] Terminal: Collects audio during meetings in real time via the microphone.
[0106] Step 2:
[0107] Terminal: Digitizes the collected audio data and sends it to the server.
[0108] Step 3:
[0109] Server: Inputs received audio data into a speech recognition model and converts it into text.
[0110] Step 4:
[0111] Server: Stores and filters the text-based data.
[0112] Step 5:
[0113] Server: Sends filtered text to the terminal.
[0114] Step 6:
[0115] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[0116] Step 7:
[0117] User: Review the displayed text and edit or save it as needed.
[0118] Notification of the meaning of unfamiliar words
[0119] Step 1:
[0120] User: Enter text into Excel sheets or documents.
[0121] Step 2:
[0122] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[0123] Step 3:
[0124] Terminal: Sends detected words and phrases to the server.
[0125] Step 4:
[0126] Server: Searches for the meaning of words and phrases, as well as related information.
[0127] Step 5:
[0128] Server: Create search results as a pop-up notification.
[0129] Step 6:
[0130] Device: Displays a pop-up notification to the user.
[0131] Step 7:
[0132] User: Review the displayed explanation and access any further information you need.
[0133] Suggestion for an automated reply
[0134] Step 1:
[0135] Terminal: Notifies the user that a new email has been received.
[0136] Step 2:
[0137] Terminal: Sends the content of received emails to the server.
[0138] Step 3:
[0139] Server: Analyzes the content of the email and generates an initial reply.
[0140] Step 4:
[0141] Server: Sends the generated initial reply message to the terminal.
[0142] Step 5:
[0143] Terminal: Displays the initial reply as a draft.
[0144] Step 6:
[0145] User: Review the draft, make any necessary corrections, and then submit.
[0146] Function suggestions for Excel and spreadsheets
[0147] Step 1:
[0148] User: Enter data into Excel or a spreadsheet.
[0149] Step 2:
[0150] Terminal: Analyzes the input data and requests function suggestions from the server.
[0151] Step 3:
[0152] Server: Selects the most appropriate function or script based on the received data.
[0153] Step 4:
[0154] Server: Generates the selected functions and scripts as suggestions and sends them to the terminal.
[0155] Step 5:
[0156] Terminal: Automatically inserts the suggested function into the cell.
[0157] Step 6:
[0158] User: Review the proposed function and set the required data range and conditions.
[0159] Task organization and reminder function
[0160] Step 1:
[0161] User: Enter tasks into a task management app or scheduler.
[0162] Step 2:
[0163] Terminal: Sends entered task data to the server in real time.
[0164] Step 3:
[0165] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[0166] Step 4:
[0167] Server: Generates reminder notifications for identified tasks.
[0168] Step 5:
[0169] Server: Sends a reminder notification to the device.
[0170] Step 6:
[0171] Device: Displays reminder notifications to the user at the appropriate time.
[0172] Step 7:
[0173] User: Check the reminder notification and take action on the task.
[0174] (Example 1)
[0175] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0176] In modern work and daily life, there is a demand for efficiently managing and executing a wide range of tasks. In particular, there is a lack of tools to streamline tasks such as recording meeting content, clarifying the meaning of unfamiliar words, responding quickly to emails, analyzing complex data and proposing functions, and task reminders. Furthermore, users face the problem of having to spend a significant amount of time and effort performing these tasks individually.
[0177] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0178] In this invention, the server includes means for analyzing audio data and converting it into text, means for searching for the meaning and related information of detected words and phrases, means for generating a preliminary reply and displaying it as a draft on the terminal, means for analyzing data and generating appropriate functions and scripts, and means for analyzing task information and creating reminder notifications. This makes it possible to efficiently manage a wide range of tasks and significantly reduce the user's working time and effort.
[0179] "Audio data" refers to data that represents audio in a digital format.
[0180] A "server" is a computer system that processes and stores data over a network.
[0181] A "device" refers to a computer, smartphone, tablet, or other device used by a user.
[0182] A "user" is an individual who operates this system to perform work-related or daily life tasks.
[0183] A "microphone" is a device that collects sound and converts it into an electrical signal.
[0184] "Analysis" is the process of breaking down and interpreting data to reveal its contents.
[0185] "Text" is a collection of information expressed in written form.
[0186] "Real-time" refers to processing that occurs almost simultaneously without delay.
[0187] "Editing" is the process of modifying or changing text or other data.
[0188] "Saving" refers to the act of storing data in a storage device.
[0189] An "unknown word" is a word or phrase whose meaning is unknown to the user or system.
[0190] "Searching" is the process of finding specific information.
[0191] A "pop-up notification" is a notification window that temporarily appears on the screen.
[0192] A "first reply" is the initial draft of a response in communication such as email.
[0193] A "draft" is a preliminary version of a document created before it becomes a formal document.
[0194] A "function" is a mathematical or programming operation that takes a number or string as input and outputs a specific result.
[0195] A "script" is a series of commands or instructions that automatically execute a specific process.
[0196] A "task" is a specific activity or action in one's work or daily life.
[0197] A "reminder notification" is a notification that reminds you of the completion or deadline of a specific task.
[0198] A "generative AI model" is an artificial intelligence model that uses machine learning techniques to generate new data.
[0199] This invention relates to a resident generation AI system for users to efficiently manage and perform a wide range of tasks in their work and daily lives. This system provides real-time text conversion of voice data, notification of the meaning of unknown words, suggestions for automatic replies, suggestions for functions and scripts, and task reminder functions. Specific embodiments of this invention are described below.
[0200] System Configuration
[0201] Real-time transcription of meeting content
[0202] 1. Collection and transmission of audio data
[0203] Terminal: Collects audio in real time via the microphone during meetings and conversations. For example, use an audio capture library (e.g., PortAudio) on the terminal's operating system.
[0204] Terminal: Digitizes the collected audio data, performs encoding, and then sends it to the server using the HTTPS protocol.
[0205] 2. Analysis of audio data and text conversion
[0206] Server: Inputs the received audio data into a speech recognition API such as Google Cloud Speech-to-Text and converts it into text data.
[0207] 3. Temporary storage and real-time display of text data
[0208] Server: Temporarily stores the converted text data in a database (e.g., MySQL®).
[0209] Terminal: Displays filtered text data in a dedicated window in real time. For example, it updates the page using JavaScript® and WebSocket.
[0210] User: Review the displayed text and edit or save it as needed. The edited text will be saved to your local disk or cloud storage (e.g., Google Drive).
[0211] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server using Google Cloud Speech-to-Text. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[0212] Notification of the meaning of unfamiliar words
[0213] 1. Text input and analysis
[0214] User: Enter text into Excel or Google Docs.
[0215] Terminal: Analyzes the input text in real time and detects words that do not exist in the dictionary database (e.g., Oxford Dictionary API).
[0216] 2. Word meaning search and notification
[0217] Server: Uses the received word to call the Oxford Dictionaries API and retrieve its meaning and related information.
[0218] Device: Display search results as a pop-up notification. For example, using DOM manipulation and JavaScript.
[0219] User: Check the displayed explanation and obtain the necessary information.
[0220] Specific example: If a user tries to use the SUM function in Excel and doesn't understand its meaning, simply selecting the function name will cause the terminal to display a pop-up based on information from the Oxford Dictionaries. This allows the user to understand the meaning of the function without interrupting their work.
[0221] Suggestion for an automated reply
[0222] 1. Receiving and notifying emails
[0223] Terminal: Notifies the user that a new email has been received and displays its contents in the email client (e.g., Microsoft® Outlook).
[0224] 2. Content analysis and generation of initial reply
[0225] Server: Analyzes the email content using a generation AI model (e.g., OpenAI®, GPT-4®) and generates an initial reply.
[0226] 3. Draft display and confirmation
[0227] Terminal: Displays the generated initial reply as a draft.
[0228] User: Review the draft, make any necessary corrections, and send a reply.
[0229] Specific example: When a user receives a new email, the server analyzes its contents and generates an initial reply message such as, "Can you make any necessary adjustments?" The user can then review it, make any necessary corrections, and quickly send a reply.
[0230] Function suggestions for Excel and spreadsheets
[0231] 1. Data Input and Analysis
[0232] User: Enters data into Microsoft Excel (registered trademark) or Google Sheets.
[0233] Terminal: Analyzes input data in real time. Using the Python pandas library is recommended.
[0234] 2. Function suggestions and input assistance
[0235] Server: Selects the most appropriate function based on the data content and generates a proposal using an AI model (e.g., OpenAI).
[0236] Terminal: Automatically inserts the suggested function into the cell.
[0237] User: Review the proposed function and set the data range and conditions.
[0238] Specific example: When a user aggregates sales data in a spreadsheet, a function like "=SUM(A2:A10)" is suggested and automatically inserted. The user can then easily check the data range and perform the calculation.
[0239] Task organization and reminder function
[0240] 1. Enter and submit task information
[0241] User: Enter task information into Microsoft To-Do or Google Calendar.
[0242] Terminal: Sends entered task information to the server in real time. Uses the HTTPS protocol.
[0243] 2. Task analysis and creation of reminder notifications
[0244] Server: Analyzes task information, extracts tasks with approaching deadlines, and creates reminder notifications.
[0245] 3. Displaying reminder notifications
[0246] Device: Displays reminder notifications to the user at the appropriate time.
[0247] User: Check reminder notifications and manage tasks.
[0248] Specific example: A user uses a task management app, and as the deadline for a task approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user receives this notification and can manage their tasks appropriately.
[0249] Examples of prompts for generative AI models
[0250] 1. "I want to collect the audio from the meeting and display it as text in real time. Please convert this audio to text."
[0251] 2. "I want to understand the meaning of functions used in Excel, so please explain the SUM function."
[0252] 3. "You have received a new email. Please generate an initial reply to this email immediately."
[0253] 4. "I'm entering sales data in Google Sheets, so please suggest an appropriate function."
[0254] 5. "I have set up reminder notifications in Google Calendar. The deadline for this task is approaching, so please create a reminder notification."
[0255] This invention's system allows users to efficiently carry out various tasks in their work and daily lives. It reduces the time and effort required for information retrieval, data entry, task management, etc., and provides a more efficient work environment.
[0256] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0257] Real-time transcription of meeting content
[0258] Step 1: Collect audio data
[0259] Terminal: Collects audio data via the microphone during meetings and conversations. This audio data is an analog signal and is sent from the microphone to the terminal's audio input port.
[0260] Input: Analog audio signal collected via microphone.
[0261] Output: An analog audio signal is sent to the terminal.
[0262] Step 2: Digitize and transmit audio data
[0263] Terminal: Performs an ADC (analog-to-digital converter) to convert analog audio signals into digital data. After conversion, encodes the digital audio data (e.g., MP3 or WAV) and sends it to the server using the HTTPS protocol.
[0264] Input: Analog audio signal.
[0265] Output: Digital audio data is sent to the server.
[0266] Step 3: Analyzing the audio data
[0267] Server: Inputs the received digital audio data into a speech recognition API (e.g., Google Cloud Speech-to-Text) for analysis. The API converts the audio data into text data.
[0268] Input: Digital audio data.
[0269] Output: Convert to text data.
[0270] Step 4: Save text data
[0271] Server: Temporarily stores the converted text data in a database (e.g., MySQL).
[0272] Input: Converted text data.
[0273] Output: Temporarily saved text data.
[0274] Step 5: Filter the text data.
[0275] Server: Extracts important keywords and phrases from stored text data and performs filtering to remove unnecessary noise. For example, natural language processing (NLP) techniques are used to extract content appropriate to the meeting context.
[0276] Input: Temporarily saved text data.
[0277] Output: Filtered text data.
[0278] Step 6: Real-time Display of Text
[0279] Terminal: The filtered text data is displayed in real time in a dedicated window. For example, the page is updated using JavaScript and WebSocket.
[0280] Input: Filtered text data.
[0281] Output: Text displayed in real time.
[0282] Step 7: Editing and Saving of Text
[0283] User: Check the displayed text and edit and save it as necessary. The edited text is saved on the local disk or cloud storage (e.g., Google Drive).
[0284] Input: Text displayed in real time.
[0285] Output: Edited and saved text data.
[0286] Notification of the Meaning of Unknown Words
[0287] Step 1: Input of Text
[0288] User: Enter text in Excel or Google Docs.
[0289] Input: Text entered by the user.
[0290] Output: The text is input into the terminal.
[0291] Step 2: Detection of Unknown Words
[0292] Terminal: Analyze the input text in real time and detect words that do not exist in the dictionary database (e.g., Oxford Dictionary API).
[0293] Input: The input text.
[0294] Output: A list of unknown words.
[0295] Step 3: Sending words
[0296] Terminal: Send the unknown words to the server. For example, use an Ajax request.
[0297] Input: A list of unknown words.
[0298] Output: The unknown words sent to the server.
[0299] Step 4: Searching and analyzing meanings
[0300] Server: Use the dictionary API to search for the meanings and related information of the sent words and analyze the related information.
[0301] Input: The unknown words sent to the server.
[0302] Output: Meanings and related information.
[0303] Step 5: Notifying meanings
[0304] Terminal: Display the analyzed meanings and related information as a popup window. Specifically, use DOM manipulation and JavaScript.
[0305] Input: Meanings and related information.
[0306] Output: The meanings displayed in the popup notification.
[0307] Step 6: Confirming content
[0308] User: Check the displayed explanation and obtain the necessary information.
[0309] Input: Meaning of the pop-up notification.
[0310] Output: Information obtained by the user.
[0311] Suggestion for an automated reply
[0312] Step 1: Receiving and receiving emails
[0313] Terminal: Notifies the user that a new email has been received and displays its contents in the email client (e.g., Microsoft Outlook).
[0314] Input: Received email.
[0315] Output: Email notification displayed in the email client.
[0316] Step 2: Analyze the content of the email
[0317] Server: Sends the body of the received email to a generating AI model (e.g., OpenAI GPT-4) for analysis.
[0318] Input: The body of the received email.
[0319] Output: Analyzed email content.
[0320] Step 3: Generating the initial reply
[0321] Server: Based on the analysis results, it generates an initial reply message and applies its content to a template.
[0322] Input: Analyzed email content.
[0323] Output: Initial reply.
[0324] Step 4: Draft Display
[0325] Terminal: Displays the generated initial reply as a draft.
[0326] Input: Initial reply message.
[0327] Output: Reply displayed as a draft.
[0328] Step 5: Review and revise the draft
[0329] User: Review the draft and make any necessary corrections.
[0330] Input: Reply displayed as a draft.
[0331] Output: Revised reply text.
[0332] Step 6: Send a reply
[0333] User: Send the revised reply.
[0334] Input: Revised reply text.
[0335] Output: Sent reply email.
[0336] Function suggestions for Excel and spreadsheets
[0337] Step 1: Enter data
[0338] User: Enter data into Microsoft Excel or Google Sheets.
[0339] Input: Data entered by the user.
[0340] Output: Data entered into the terminal.
[0341] Step 2: Analysis of the month
[0342] Terminal: Analyzes input data in real time. Uses Python's pandas library, etc.
[0343] Input: The entered data.
[0344] Output: Analysis results.
[0345] Step 3: Propose a function
[0346] Server: Based on the data analysis results, it selects the optimal function and generates proposals using a generative AI model (e.g., OpenAI).
[0347] Input: Analysis results.
[0348] Output: Proposed function.
[0349] Step 4: Function display and input assistance
[0350] Terminal: Automatically inserts the suggested function into the cell.
[0351] Input: Proposed function.
[0352] Output: The function inserted into the cell.
[0353] Step 5: Verify and configure the function
[0354] User: Review the proposed function and set the required data range and conditions.
[0355] Input: The function inserted into the cell.
[0356] Output: The configured function.
[0357] Task organization and reminder function
[0358] Step 1: Enter task information
[0359] User: Enter task information into Microsoft To-Do or Google Calendar.
[0360] Input: Task information entered by the user.
[0361] Output: Task information entered into the terminal.
[0362] Step 2: Submit task data
[0363] Terminal: Sends entered task information to the server in real time. Uses the HTTPS protocol.
[0364] Input: Task information.
[0365] Output: Task information sent to the server.
[0366] Step 3: Task Analysis
[0367] Server: Analyzes task data and extracts tasks with approaching deadlines.
[0368] Input: Submitted task information.
[0369] Output: Extracted tasks with approaching deadlines.
[0370] Step 4: Create a reminder notification
[0371] Server: Creates reminder notifications for tasks with approaching deadlines.
[0372] Input: Extracted tasks with approaching deadlines.
[0373] Output: Reminder notification.
[0374] Step 5: Displaying reminder notifications
[0375] Terminal: Displays the created reminder notification to the user at the appropriate time.
[0376] Input: Reminder notification.
[0377] Output: Reminder notification displayed to the user.
[0378] Step 6: Task Management
[0379] User: Check reminder notifications and manage tasks.
[0380] Input: The displayed reminder notification.
[0381] Output: Managed tasks.
[0382] The above steps establish the processing flow of this system. This allows users to efficiently manage and perform various tasks and daily routines.
[0383] (Application Example 1)
[0384] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0385] In modern security operations, there is a demand for efficient information collection and management, as well as rapid response. However, manual reporting during patrols is time-consuming and laborious, and prone to information omissions and communication errors. Furthermore, daily tasks such as task reminders and initial email replies can add to the workload, reducing overall efficiency. In this context, there is a growing need for systems that convert speech to text in real time, automatically generate responses, and provide task reminders.
[0386] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0387] In this invention, the server includes means for collecting audio data through the microphone of a terminal used by the user, means for transmitting the collected audio data to the server, means for the server to analyze the audio data and convert it into text, means for generating an appropriate response sentence using a generation AI model based on the converted text, means for displaying the generated response sentence on the terminal in real time, means for transmitting task data registered by the user to the server, means for the server to analyze the task data and create a reminder notification for tasks with approaching deadlines, and means for displaying the reminder notification on the terminal. This enables real-time information gathering and response in security operations, thereby improving operational efficiency.
[0388] "User" refers to an individual or group that uses the system.
[0389] "Terminal" refers to a digital device used by a user (e.g., a smartphone or PC).
[0390] A "microphone" refers to a device that converts sound into electrical signals.
[0391] "Audio data" refers to the digital signal of sound collected through a microphone.
[0392] A "server" refers to a computer system that receives audio data, analyzes it, and converts it into text.
[0393] A "generative AI model" refers to artificial intelligence that generates appropriate responses or suggestions based on input text data.
[0394] "Text" refers to written information converted from audio data.
[0395] "Response sentence" refers to a dialogue-style reply sentence generated by a generative AI model.
[0396] "Task data" refers to information about tasks and schedules registered by the user.
[0397] A "reminder notification" refers to a notification that informs the user of tasks that are nearing their deadline.
[0398] This invention is a system for streamlining information gathering and task management in security operations. The system includes means for collecting voice data through the microphone of a terminal used by the user and transmitting the collected voice data to a server. The server analyzes the voice data, converts it to text, generates appropriate response sentences using a generative AI model, and displays them on the terminal in real time. Furthermore, the system transmits task data registered by the user to the server, which analyzes the task data, creates reminder notifications for tasks with approaching deadlines, and displays these on the terminal.
[0399] The server uses the Google Cloud Speech-to-Text API to convert speech data into text and generates response sentences based on the text data using a generative AI model (e.g., GPT-4). The terminal displays the transcribed meeting content and generated response sentences to the user in real time. This allows users to quickly and accurately obtain information and manage tasks during security work.
[0400] For example, if a night shift security staff member uses a mobile device to make a voice report during patrol, the voice is automatically converted to text and sent to the administrator in real time. Furthermore, reminder notifications are displayed based on a security checklist, helping to prevent missed tasks. This is expected to improve work efficiency and reduce errors.
[0401] The system of the present invention is specifically implemented through the following program processing. The program uses the Google Cloud Speech-to-Text API to convert speech data into text and generates response sentences using a generative AI model. Furthermore, a task reminder function is implemented using a scheduling library.
[0402] For example, the following prompt statements are possible:
[0403] "You are the designer of an application that helps streamline security services. It will transcribe voice data into text in real time and use AI to provide users with appropriate replies and task notifications. Please design an application with the following features:
[0404] Audio data collection and real-time text conversion
[0405] Generating appropriate automated replies from text data
[0406] "Recurring task reminder notifications"
[0407] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0408] Step 1:
[0409] The device collects audio data. The device collects the voice spoken by the user through the microphone in real time and saves it as data. In this step, the input is the user's voice, and the output is digitized audio data.
[0410] Step 2:
[0411] The terminal sends the collected audio data to the server. The collected audio data is uploaded to the server in real time. The input for this step is digitized audio data, and the output is audio data stored on the server.
[0412] Step 3:
[0413] The server analyzes the audio data and converts it to text. The server uses the Google Cloud Speech-to-Text API to convert the audio data to text data. The input for this step is audio data, and the output is text data. As a concrete example of how the audio analysis process works, an audio file is input to the API, and the result is returned as text.
[0414] Step 4:
[0415] The server generates an appropriate response sentence using a generative AI model based on the converted text. The generative AI model (e.g., GPT-4) is input with text data to generate a response sentence. The input for this step is text data, and the output is the generated response sentence. The generative AI model analyzes the text data to produce the most appropriate response or suggestion.
[0416] Step 5:
[0417] The server sends the generated response message to the terminal in real time. The generated response message is sent to the terminal and displayed to the user immediately. The input for this step is the response message, and the output is the response message displayed on the terminal.
[0418] Step 6:
[0419] The user registers task data on their device, and the device sends that data to the server. When a user registers a task using a task management app, that data is sent to the server in real time. In this step, the input is the task data entered by the user, and the output is the task data stored on the server.
[0420] Step 7:
[0421] The server analyzes task data and creates reminder notifications for tasks with approaching deadlines. It analyzes task data, extracts tasks with imminent deadlines, and generates reminder notifications. The input for this step is task data, and the output is the generated reminder notifications.
[0422] Step 8:
[0423] The server sends a reminder notification to the device, and the device displays it to the user. The server sends a reminder notification to the device and notifies the user. The input for this step is the reminder notification, and the output is the reminder notification displayed on the device.
[0424] The above processing steps enable real-time text conversion of audio data, generation of response statements, task management, and reminder notifications, thereby streamlining security operations.
[0425] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0426] This invention relates to a resident generative AI system for streamlining information retrieval, text conversion, task management, email correspondence, function input, and user emotion recognition and response in users' work and daily lives. This invention comprises the following components and their operation.
[0427] System Configuration
[0428] Real-time transcription of meeting content
[0429] Collection and transmission of audio data
[0430] Terminal: Collects audio during meetings in real time via the microphone.
[0431] Terminal: Digitizes the collected audio data and sends it to the server.
[0432] Audio data analysis and text conversion
[0433] Server: Inputs received audio data into a speech recognition model and converts it into text.
[0434] Server: Stores and filters the text-based data.
[0435] Server: Sends filtered text to the terminal.
[0436] Display text
[0437] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[0438] User: Review the displayed text and edit or save it as needed.
[0439] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[0440] Notification of the meaning of unfamiliar words
[0441] Text input and analysis
[0442] User: Enter text into Excel sheets or documents.
[0443] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[0444] Terminal: Sends detected words and phrases to the server.
[0445] Word meaning search
[0446] Server: Searches for the meaning of words and phrases, as well as related information.
[0447] Server: Create search results as a pop-up notification.
[0448] Meaning notification
[0449] Device: Displays a pop-up notification to the user.
[0450] User: Review the displayed explanation and access any further information you need.
[0451] Specific example: If a user doesn't understand the meaning of a particular function in Excel, selecting that function will display its meaning in a pop-up on the device. This allows the user to understand the meaning without interrupting their work.
[0452] Suggestion for an automated reply
[0453] Receiving and analyzing emails
[0454] Terminal: Notifies the user that a new email has been received.
[0455] Terminal: Sends the content of received emails to the server.
[0456] Reply generation and display
[0457] Server: Analyzes the content of the email and generates an initial reply.
[0458] Server: Sends the generated initial reply message to the terminal.
[0459] Terminal: Displays the initial reply as a draft.
[0460] User: Review the draft, make any necessary corrections, and then send a reply.
[0461] Specific example: When a user receives an email, the server analyzes its contents and generates an initial reply such as, "Could you please schedule a meeting on the following dates?" The user can then review and send the email, allowing for a quick response.
[0462] Function suggestions for Excel and spreadsheets
[0463] Text input and analysis
[0464] User: Enter data into Excel or a spreadsheet.
[0465] Terminal: Analyzes the input data and requests function suggestions from the server.
[0466] Function proposal generation and input assistance
[0467] Server: Selects the most appropriate function or script based on the received data.
[0468] Server: Generates the selected functions and scripts as suggestions and sends them to the terminal.
[0469] Terminal: Automatically inserts the suggested function into the cell.
[0470] User: Review the proposed function and set the required data range and conditions.
[0471] Specific example: When a user aggregates sales data, the "SUM function" is suggested, and the terminal automatically inserts the function. The user can perform the aggregation simply by specifying the data range.
[0472] Task organization and reminder function
[0473] Task data collection and transmission
[0474] User: Enter tasks into a task management app or scheduler.
[0475] Terminal: Sends entered task data to the server in real time.
[0476] Task analysis and reminders
[0477] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[0478] Server: Generates reminder notifications for identified tasks.
[0479] Server: Sends a reminder notification to the device.
[0480] Display reminder notifications
[0481] Device: Displays reminder notifications to the user at the appropriate time.
[0482] User: Check the reminder notification and take action on the task.
[0483] Specific example: When a user uses a task management app and the deadline approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user can then see the reminder and efficiently work on the task.
[0484] Additional configuration for the emotion engine
[0485] Recognition and response to emotions
[0486] Collection and analysis of voice and input data
[0487] Device: Collects voice and text input in real time.
[0488] Terminal: Sends collected data to the server.
[0489] Recognition of emotions
[0490] Server: Uses an emotion engine to analyze the user's emotional state from collected audio and text data.
[0491] Server: Classifies the user's emotional state based on the analysis results and determines appropriate feedback and responses.
[0492] Feedback and implementation of countermeasures
[0493] Server: Generates feedback and responses based on the user's emotional state and sends them to the terminal.
[0494] Terminal: Displays generated feedback and responses to the user.
[0495] Example 1: When a user receives an emotionally charged email, the emotion engine detects the user's stress level and suggests softening the tone of the reply. The device displays the suggestion, and the user can send it as is.
[0496] Example 2: If a user is working with a spreadsheet and the emotion engine detects the user's confusion or stress, the device will display appropriate guidance and support to help the user.
[0497] Specific example 3: In user task management, if the emotion engine detects user fatigue or anxiety, the server adjusts task priorities and sends a reminder notification to the device. The user can then check the notification and efficiently proceed with their tasks.
[0498] By combining the system of the present invention with an emotion engine, it becomes possible to recognize the user's emotional state and provide appropriate feedback and support. This allows users to work more comfortably and efficiently in their professional and daily lives.
[0499] The following describes the processing flow.
[0500] Additional configuration for the emotion engine
[0501] Recognition and response to emotions
[0502] Collection and analysis of voice and input data
[0503] Step 1:
[0504] Terminal: Collects audio from the user's voice through the microphone.
[0505] Step 2:
[0506] Terminal: Collects text data entered by the user.
[0507] Step 3:
[0508] Terminal: Sends collected audio and text data to the server.
[0509] Recognition of emotions
[0510] Step 1:
[0511] Server: Analyzes the collected audio data and determines the user's emotional state from the audio data.
[0512] Step 2:
[0513] Server: Analyzes collected text data and uses that data to analyze the user's emotional state.
[0514] Step 3:
[0515] Server: Integrates the analysis results of voice and text data to classify the user's overall emotional state.
[0516] Feedback and implementation of countermeasures
[0517] Step 1:
[0518] Server: Determines appropriate feedback and support based on emotional state.
[0519] Step 2:
[0520] Server: Generates the determined feedback and support details and sends them to the terminal.
[0521] Step 3:
[0522] Terminal: Displays generated feedback and support details to the user.
[0523] Specific example
[0524] Example 1: Emotional email responses
[0525] Step 1:
[0526] Terminal: Collects the content of emails received by the user and sends it to the server.
[0527] Step 2:
[0528] Server: Analyzes email content and uses an emotion engine to analyze the user's emotional state.
[0529] Step 3:
[0530] Server: Based on the analysis results, it generates a reply message to alleviate user stress.
[0531] Step 4:
[0532] Terminal: Displays the generated reply as a draft to the user.
[0533] Step 5:
[0534] User: Review the draft, make any necessary corrections, and then send a reply.
[0535] Example 2: Spreadsheet operation support
[0536] Step 1:
[0537] Terminal: Collects data entered by the user in a spreadsheet and sends it to the server.
[0538] Step 2:
[0539] Server: Uses an emotion engine to analyze the confusion and stress the user is experiencing.
[0540] Step 3:
[0541] Server: Generates appropriate guides and support content based on the analysis results.
[0542] Step 4:
[0543] Terminal: Displays generated guides and support information to the user.
[0544] Specific example 3: Adjusting task management and reminders
[0545] Step 1:
[0546] Terminal: Collects task data entered by the user into the task management app and sends it to the server.
[0547] Step 2:
[0548] Server: Uses an emotion engine to analyze user fatigue and anxiety.
[0549] Step 3:
[0550] Server: Based on the analysis results, adjust task priorities and reminder timings.
[0551] Step 4:
[0552] Server: Generates a customized reminder notification and sends it to the device.
[0553] Step 5:
[0554] Terminal: Displays the generated reminder notification to the user at the appropriate time.
[0555] Step 6:
[0556] User: Check the displayed reminder notification and take action on the task.
[0557] Thus, by combining an emotion engine, the system of the present invention can recognize the user's emotional state in real time and provide appropriate feedback and support, thereby improving efficiency in work and daily life.
[0558] (Example 2)
[0559] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0560] Modern information processing devices and various user systems require the efficient, real-time processing of a wide variety of tasks. However, centralized systems for quickly and accurately performing tasks such as converting audio data to text, searching for the meaning of unknown words and phrases, and automatically replying to emails are still not adequately developed. The lack of such systems can lead to decreased work efficiency and increased stress for users.
[0561] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing audio data and converting it into text using a speech recognition model, means for using the internet to search for the meaning of words and phrases and related information, and means for analyzing the content of emails and generating a primary reply using a generation AI model. This enables users to quickly record meeting content, instantly understand the meaning of unfamiliar words, and efficiently reply to emails.
[0562] "Audio data" refers to information that represents audio from meetings, conversations, etc., in a digital format.
[0563] "Digitalization" is the process of converting analog signals into digital signals.
[0564] A "server" is a computer that provides services and data to multiple clients over a network.
[0565] A "speech recognition model" is a machine learning model used to convert speech data into text.
[0566] "Text conversion" is the process of converting data such as audio or handwritten notes into text data.
[0567] "Filtering" is the process of applying specific conditions to data to extract only the necessary information.
[0568] An "information processing device" is an electronic computer that has the functions of inputting, processing, and outputting data.
[0569] An "unknown word or phrase" is a word or part of a sentence that the user cannot understand or find unclear.
[0570] A "pop-up notification" is a notification window that temporarily appears on the user's screen.
[0571] "Email" refers to messages that are sent and received electronically via the internet.
[0572] A "generative AI model" is a machine learning model used to generate text and data using artificial intelligence.
[0573] A "primary reply" is the initial response suggested in an email reply.
[0574] A "draft" is a preliminary version of a document or email before any revisions are made.
[0575] This invention relates to a resident generative AI system for streamlining information retrieval, text conversion, task management, email correspondence, function input, and user emotion recognition and response in users' work and daily lives. This invention comprises the following components and their operation.
[0576] Real-time transcription of meeting content
[0577] This system collects audio data in real time during meetings through the microphone of the user's information processing device, digitizes it, and sends it to a server. The server converts the received audio data into text using a speech recognition model (e.g., Google Cloud Speech-to-Text API). The transcribed data is stored in a database, where important keywords are filtered. The filtered text is sent to the information processing device in real time and displayed in a dedicated window.
[0578] Specific example:
[0579] When a user is conducting an online meeting, the information processing device collects audio data through its microphone and sends the digitized data to a server. The server uses a speech recognition model to transcribe the data into text, filters it, and then sends it back to the information processing device, allowing the meeting content to be displayed as text in real time.
[0580] Example of a prompt:
[0581] "Please explain, step by step, the process of the system that transcribes meeting content into text in real time."
[0582] Notification of the meaning of unfamiliar words
[0583] The system analyzes the text entered by the user into the information processing device in real time to detect unknown words and phrases. These detected words and phrases are sent to a server. The server searches for the meaning and related information of these words and phrases using the internet (e.g., Wikipedia API, Oxford Dictionaries API) and displays it on the information processing device as a pop-up notification.
[0584] Specific example:
[0585] When a user enters a specific function into an Excel sheet, it is detected and sent to the server. The server uses the corresponding API to look up the meaning of the function and displays it on the information processing device as a pop-up notification.
[0586] Example of a prompt:
[0587] "Please explain, step by step, the process of a system that automatically notifies you of the meaning of unknown words in Excel."
[0588] Suggestion for an automated reply
[0589] The server analyzes the content of the received email and generates a preliminary reply using a generative AI model (e.g., OpenAI GPT-3®). The generated preliminary reply is displayed as a draft on the information processing device, and the user reviews it, makes corrections, and then sends it.
[0590] Specific example:
[0591] When a user receives an email, the server analyzes its contents and generates a preliminary reply such as, "Would you be able to schedule a meeting on the following dates?" The user can then review, revise, and send the reply, which is displayed as a draft.
[0592] Example of a prompt:
[0593] "Please explain the process of the system that generates automated reply messages, step by step."
[0594] Task organization and reminder function
[0595] When a user enters a task into a task management app or scheduler, that task data is sent to the server in real time. The server analyzes the task data, identifies unattended tasks with deadlines, and generates reminder notifications. These reminder notifications are sent to an information processing device and displayed to the user at the appropriate time.
[0596] Specific example:
[0597] When a user uses a task management app and the deadline approaches, the information processing device notifies them with a message such as, "Please complete this task by 3 PM tomorrow." The user can then use the reminder to efficiently complete the task.
[0598] Example of a prompt:
[0599] "Please describe the process of the system that sends task reminder notifications, step by step."
[0600] Recognition and response to emotions
[0601] This system incorporates an emotion engine that analyzes the user's emotional state using voice and text data collected from the user. This allows the system to classify the user's emotional state and determine appropriate feedback and responses. The generated feedback and responses are displayed on the information processing device.
[0602] Specific example:
[0603] When a user receives an emotionally charged email, the emotion engine detects the user's stress and suggests softening the tone of the reply. If the emotion engine detects confusion or stress while a user is working with a spreadsheet, the device displays appropriate guidance and support. In task management, if the emotion engine detects user fatigue or anxiety, the server adjusts task priorities and sends reminder notifications to the information processing device.
[0604] Example of a prompt:
[0605] "Please describe, step by step, how the system recognizes and appropriately responds to user emotions."
[0606] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0607] Real-time transcription of meeting content
[0608] Step 1:
[0609] Voice collection
[0610] Device: Collects audio during meetings in real time using a microphone.
[0611] Input: Analog audio from a meeting.
[0612] Output: Analog audio data.
[0613] Specific operation: By pressing the "Start Recording" button, the device activates the microphone and begins collecting audio data.
[0614] Step 2:
[0615] Audio digitization
[0616] Terminal: Digitizes the collected analog audio data.
[0617] Input: Analog audio data.
[0618] Output: Digital audio data.
[0619] Specific operation: Digitize the audio at a sampling rate of 44.1kHz and convert it to PCM format.
[0620] Step 3:
[0621] Sending audio data
[0622] Terminal: Sends digitized audio data to the server.
[0623] Input: Digital audio data.
[0624] Output: Transmission of digital audio data to the server.
[0625] Specific operation: Digitized audio data is sent to the server using a real-time streaming protocol (e.g., WebSocket).
[0626] Step 4:
[0627] Receiving audio data
[0628] Server: The server receives audio data in real time.
[0629] Input: Digital audio data transmitted from the device.
[0630] Output: Audio data stored in the buffer.
[0631] Specific operation: The server continuously receives frames of audio data and stores them in a fixed buffer.
[0632] Step 5:
[0633] Speech recognition and text conversion
[0634] Server: Converts received audio data into text using a speech recognition model.
[0635] Input: Audio data stored in the buffer.
[0636] Output: Text data.
[0637] Specific operation: When the audio data buffer reaches a certain amount, the data is sent to the speech recognition API as a request, and the text result is received.
[0638] Step 6:
[0639] Text filtering and saving
[0640] Server: Stores the digitized data in a database and filters for important keywords.
[0641] Input: Text data.
[0642] Output: Filtered text data.
[0643] Specific operation: When saving text to the database, an SQL query is used to set the "importance" field, and if a specific keyword is included, it is set to "high," etc.
[0644] Step 7:
[0645] Sending filtered text
[0646] Server: Transmits text data to the information processing device in real time.
[0647] Input: Filtered text data.
[0648] Output: Sending text data to an information processing device.
[0649] Specific operation: The server sends filtered text data to the terminal via WebSocket.
[0650] Step 8:
[0651] Display text
[0652] Terminal: Displays text data in real time in a dedicated window.
[0653] Input: Filtered text data.
[0654] Output: The text displayed to the user.
[0655] Specific operation: Uses Javascript and other front-end technologies to manipulate the DOM within the window and render text.
[0656] Step 9:
[0657] Edit and save text
[0658] User: Review the displayed text, edit it as needed, and save it.
[0659] Input: The displayed text.
[0660] Output: Saved text data.
[0661] Specific operation: A text editor is provided on the editing screen, and the data is saved to local storage or cloud storage by pressing the "Save" button.
[0662] Notification of the meaning of unfamiliar words
[0663] Step 1:
[0664] Text input
[0665] User: Enter text into Excel or other documents.
[0666] Input: Text entered by the user.
[0667] Output: Input text data.
[0668] Specific action: The user enters formulas and explanations into an Excel sheet.
[0669] Step 2:
[0670] Real-time analysis
[0671] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[0672] Input: The entered text data.
[0673] Output: Unknown words or phrases detected.
[0674] Specific operation: Every time text is entered, it is analyzed using NLP (e.g., spaCy).
[0675] Step 3:
[0676] Sending to the server
[0677] Terminal: Sends detected words and phrases to the server.
[0678] Input: Unknown words or phrases detected.
[0679] Output: Data sent to the server.
[0680] Specific operation: Encode the detected data into JSON format and send it to the server via an HTTP request.
[0681] Step 4:
[0682] Search for meaning
[0683] Server: The server searches for the meaning of words and phrases using the internet.
[0684] Input: Unknown words or phrases detected.
[0685] Output: Search results data.
[0686] Specific operation: Send requests to the Wikipedia API and other dictionary APIs to retrieve meanings and related information.
[0687] Step 5:
[0688] Pop-up notification generation
[0689] Server: Generates a popup notification based on the search results.
[0690] Input: Search results data.
[0691] Output: Popup notification data.
[0692] Specific operation: Format the retrieved search results and generate a popup notification using HTML and CSS.
[0693] Step 6:
[0694] Send to device
[0695] Server: Sends a pop-up notification to the device.
[0696] Input: Pop-up notification data.
[0697] Output: Data sent to the terminal.
[0698] Specific action: The created popup notification is sent to the device as an HTTP response in JSON format.
[0699] Step 7:
[0700] Pop-up display
[0701] Device: Displays a pop-up notification on the user's screen.
[0702] Input: Pop-up notification data from the server.
[0703] Output: A pop-up notification displayed on the screen.
[0704] Specific operation: Use JavaScript to display a popup at the appropriate location and timing.
[0705] Suggestion for an automated reply
[0706] Step 1:
[0707] Email notification
[0708] Terminal: Notifies the user that a new email has been received.
[0709] Input: Received email.
[0710] Output: Notification of incoming email.
[0711] Specific action: Notify the user via a pop-up notification that a new email has been received.
[0712] Step 2:
[0713] Sending the email content
[0714] Terminal: Sends the content of received emails to the server.
[0715] Input: Received email data.
[0716] Output: Data sent to the server.
[0717] Specific operation: Convert the email content to JSON format and send it to the server via an HTTP request.
[0718] Step 3:
[0719] Generating a reply
[0720] Server: The server analyzes the email content and generates an initial reply using a generation AI model.
[0721] Input: Email content data.
[0722] Output: Initial reply.
[0723] Specific operation: Analyze the content of the email and generate a reply using a generative AI model such as OpenAI GPT-3.
[0724] Step 4:
[0725] Sending a reply
[0726] Server: Sends the generated initial reply to the terminal as a draft.
[0727] Input: Initial reply.
[0728] Output: Data sent to the terminal.
[0729] Specific operation: Convert the reply text to JSON format and send it to the terminal as an HTTP response.
[0730] Step 5:
[0731] Displaying the reply
[0732] Terminal: Displays the initial reply as a draft.
[0733] Input: Initial response data from the server.
[0734] Output: Draft below the user.
[0735] Specific operation: Use JavaScript to display a draft of the reply message in the browser.
[0736] Step 6:
[0737] User verification and correction
[0738] User: Review the draft and make any necessary corrections.
[0739] Input: Draft of the initial reply.
[0740] Output: Revised reply.
[0741] Specific operation: The user makes modifications through a text editor in the browser.
[0742] Step 7:
[0743] Corrected and sent
[0744] User: Send the revised reply.
[0745] Input: Revised reply.
[0746] Output: Sent email.
[0747] Specific action: Press the "Send" button to send the revised reply to the recipient.
[0748] The above steps demonstrate how each function is specifically implemented. Clearly defining the specific actions and data flows performed at each processing step makes it easier to understand the overall operation of the system.
[0749] (Application Example 2)
[0750] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0751] In modern brick-and-mortar stores, employees are expected to respond to customers quickly and accurately. However, employees may be unable to provide appropriate service due to lack of experience or information. Furthermore, understanding and responding appropriately to customer emotions is also difficult. A system is needed that allows employees to receive information in real time and provide the best possible response based on customer emotions.
[0752] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing sound and text information and recognizing the emotional state of the user, means for presenting feedback and responses based on the emotional state, means for optimizing dialogue with the customer using the emotional analysis results, and means for providing support information to a visual display device for quickly and appropriately answering customer questions. This enables employees in physical stores to respond to customers quickly and appropriately.
[0753] An "information processing device" is a device used to process, store, and transmit data, and includes smartphones, tablets, and computers.
[0754] "Sound" is a wave produced by vibrations in the air, and is an audio signal collected through an input device such as a microphone.
[0755] A "network" is a communication system that allows multiple information processing devices to send and receive data from one another.
[0756] "Textual information" refers to data obtained by analyzing audio or handwritten data and converting it into text format.
[0757] "Real-time" means that data is processed and displayed almost simultaneously, with virtually no delay.
[0758] "Emotional state" refers to the emotional state of a user as analyzed using an emotion engine, and includes emotions such as joy, anger, sadness, and surprise.
[0759] "Feedback" refers to information and advice provided based on a user's behavior and circumstances.
[0760] "Response" refers to appropriate actions or measures taken in response to a specific situation or condition.
[0761] A "visual display device" is a device used to visually display data, and includes smart glasses, displays, and head-mounted displays.
[0762] A "customer" is someone who intends to purchase or use a product or service.
[0763] "Support information" refers to information and data that are provided to help users perform specific tasks.
[0764] "Emotional analysis results" refer to data obtained after evaluating and classifying emotional states.
[0765] This invention provides an information processing system for improving the customer service capabilities of employees in physical stores. The embodiments for carrying out this invention are described in detail below.
[0766] The entire system consists of a visual display device used by the user (such as smart glasses), a microphone for collecting audio, a server for analyzing the data, and a network to connect them.
[0767] 1. Speech recognition and transcription
[0768] First, the user wears smart glasses, and questions from customers in the store are collected via a microphone. The visual display device acquires the audio and transmits it to a server over the network. The server analyzes the audio data and converts it into text using a speech recognition model.
[0769] 2. Recognition of emotional states
[0770] The server analyzes customer facial expression data acquired through visual display devices and cameras, and uses an emotion engine to recognize the customer's emotional state. Emotional states include feelings such as joy, anger, sadness, and surprise. This allows the user to understand the customer's feelings.
[0771] 3. Feedback and proposed response
[0772] The server provides appropriate feedback and responses based on the recognized emotional state of the customer. For example, if the customer is angry, it suggests a response in a calm tone. If the customer is confused, it provides detailed explanations and guidance.
[0773] 4. Provision of support information
[0774] In response to customer inquiries, the server generates appropriate answers, which are displayed in real time on the smart glasses. This allows users to instantly answer customer questions. Other relevant information and product details are also presented as needed.
[0775] These features enable employees to respond to customers quickly and accurately, leading to improved customer satisfaction.
[0776] Examples
[0777] One day, if a customer in a physical store asks, "How much does this item cost?", the microphone in the smart glasses captures the voice and transmits it to a server via the network. The server converts the voice into text and also recognizes the customer's emotions from their facial expressions. Based on the emotional state, the server provides a calm response, such as, "This item costs 299 yen." This information is displayed in real time on the smart glasses' screen, allowing employees to respond to customers quickly.
[0778] Example of a prompt
[0779] "A customer has asked a question about a product. Based on that question and the subsequent conversation, please have the AI generate an appropriate answer. Also, please suggest responses that take into account the customer's emotional state."
[0780] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0781] Step 1:
[0782] The user wears smart glasses and receives customer questions via voice. A microphone built into the smart glasses collects the voice data. This voice data is input into the user's device (smart glasses).
[0783] Step 2:
[0784] The terminal digitizes the collected audio data and transmits it to a server via the network. Here, the data format is converted, and the audio data is input to the server.
[0785] Step 3:
[0786] The server inputs the received audio data into a speech recognition model (e.g., Google Speech-to-Text API) and converts it into text information. In this process, the audio data is processed into text data, and the result is output to the server.
[0787] Step 4:
[0788] The server analyzes the converted text information and extracts specific keywords (e.g., product name, price). Furthermore, it searches for supporting information (such as price information from a database) to generate the most suitable response based on that information. These search results are then output to the server.
[0789] Step 5:
[0790] The server uses the smart glasses' camera to acquire customer facial expression data and analyzes it in real time. Using an emotion recognition engine such as DeepFace, it classifies and recognizes the customer's emotional state (e.g., anger, joy). As a result, the emotional data is output to the server.
[0791] Step 6:
[0792] The server generates an action plan based on the recognized emotional state and the content of the question. For example, if the customer is angry, it will suggest a polite way to handle the situation. This information is then sent from the server to the terminal as feedback.
[0793] Step 7:
[0794] The device (smart glasses) displays feedback information (responses and suggested actions) sent from the server in real time. The user then uses this display to respond to or suggest to the customer.
[0795] Step 8:
[0796] When a user provides an appropriate response or suggestion to a customer, the terminal additionally records that information and sends it to the server as needed to improve the relevant data. This allows the system to learn and further improve its response in the future.
[0797] This processing flow is expected to significantly improve the customer service capabilities of employees in physical stores, leading to increased customer satisfaction.
[0798] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0799] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0800] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0801] [Second Embodiment]
[0802] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0803] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0804] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0805] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0806] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0807] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0808] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0809] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0810] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0811] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0812] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0813] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0814] This invention relates to a resident generation AI system for streamlining information retrieval, text conversion, task management, email correspondence, function input, and other tasks in users' work and daily lives. This invention comprises the following components and their operation.
[0815] System Configuration
[0816] Real-time transcription of meeting content
[0817] Collection and transmission of audio data
[0818] Terminal: Collects audio in real time via microphone during meetings and conversations.
[0819] Terminal: Digitizes the collected audio data and sends it to the server.
[0820] Audio data analysis and text conversion
[0821] Server: Inputs the received audio data into the speech recognition model and converts it into text.
[0822] Server: Temporarily stores the text-based data and provides the user with appropriately filtered content.
[0823] Display text
[0824] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[0825] User: Review the displayed text and edit or save it as needed.
[0826] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[0827] Notification of the meaning of unfamiliar words
[0828] Text input and analysis
[0829] User: Enter text into Excel sheets or documents.
[0830] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[0831] Word meaning search
[0832] Server: Searches for the meaning of a word or phrase selected by the user from related databases.
[0833] Server: Analyzes relevant information and extracts appropriate content.
[0834] Meaning notification
[0835] Device: Display search results as a pop-up notification.
[0836] User: Check the displayed explanation and quickly obtain the necessary information.
[0837] Specific example: If a user doesn't understand the meaning of a particular function in Excel, selecting that function will display its meaning in a pop-up on the device. This allows the user to understand the meaning without interrupting their work.
[0838] Suggestion for an automated reply
[0839] Receiving and analyzing emails
[0840] Terminal: Notifies the user that a new email has been received and displays its contents.
[0841] Server: Analyzes the content of the email and generates an appropriate initial reply.
[0842] Reply generation and display
[0843] Server: Generates an initial reply and applies it to the template.
[0844] Terminal: Displays the generated initial reply as a draft.
[0845] User: Review the draft, make any necessary corrections, and then send a reply.
[0846] Specific example: When a user receives an email, the server analyzes its contents and generates an initial reply such as, "Could you please schedule a meeting on the following dates?" The user can then review and send the email, allowing for a quick response.
[0847] Function suggestions for Excel and spreadsheets
[0848] Text input and analysis
[0849] User: Enter data into Excel or a spreadsheet.
[0850] Terminal: Analyzes the input data and suggests appropriate functions or scripts.
[0851] Function proposal generation and input assistance
[0852] Server: Based on the input, it selects the most appropriate function or script and generates it as a suggestion.
[0853] Terminal: Automatically inserts the suggested function into the cell.
[0854] User: Review the proposed function and set the required data range and conditions.
[0855] Specific example: When a user aggregates sales data, the "SUM function" is suggested, and the terminal automatically inserts the function. The user can perform the aggregation simply by specifying the data range.
[0856] Task organization and reminder function
[0857] Task data collection and transmission
[0858] User: Enter tasks into a task management app or scheduler.
[0859] Terminal: Sends entered task data to the server in real time.
[0860] Task analysis and reminders
[0861] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[0862] Server: Creates reminder notifications for tasks with approaching deadlines.
[0863] Display reminder notifications
[0864] Device: Displays reminder notifications to the user at the appropriate time.
[0865] User: Check the reminder notification and take action on the task.
[0866] Specific example: When a user uses a task management app and the deadline approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user can then see the reminder and efficiently work on the task.
[0867] The system of the present invention allows users to efficiently carry out various tasks in their work and daily lives, significantly reducing the time and effort required for information retrieval, data entry, task management, and other related activities.
[0868] The following describes the processing flow.
[0869] Real-time transcription of meeting content
[0870] Step 1:
[0871] Terminal: Collects audio during meetings in real time via the microphone.
[0872] Step 2:
[0873] Terminal: Digitizes the collected audio data and sends it to the server.
[0874] Step 3:
[0875] Server: Inputs received audio data into a speech recognition model and converts it into text.
[0876] Step 4:
[0877] Server: Stores and filters the text-based data.
[0878] Step 5:
[0879] Server: Sends filtered text to the terminal.
[0880] Step 6:
[0881] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[0882] Step 7:
[0883] User: Review the displayed text and edit or save it as needed.
[0884] Notification of the meaning of unfamiliar words
[0885] Step 1:
[0886] User: Enter text into Excel sheets or documents.
[0887] Step 2:
[0888] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[0889] Step 3:
[0890] Terminal: Sends detected words and phrases to the server.
[0891] Step 4:
[0892] Server: Searches for the meaning of words and phrases, as well as related information.
[0893] Step 5:
[0894] Server: Create search results as a pop-up notification.
[0895] Step 6:
[0896] Device: Displays a pop-up notification to the user.
[0897] Step 7:
[0898] User: Review the displayed explanation and access any further information you need.
[0899] Suggestion for an automated reply
[0900] Step 1:
[0901] Terminal: Notifies the user that a new email has been received.
[0902] Step 2:
[0903] Terminal: Sends the content of received emails to the server.
[0904] Step 3:
[0905] Server: Analyzes the content of the email and generates an initial reply.
[0906] Step 4:
[0907] Server: Sends the generated initial reply message to the terminal.
[0908] Step 5:
[0909] Terminal: Displays the initial reply as a draft.
[0910] Step 6:
[0911] User: Review the draft, make any necessary corrections, and then submit.
[0912] Function suggestions for Excel and spreadsheets
[0913] Step 1:
[0914] User: Enter data into Excel or a spreadsheet.
[0915] Step 2:
[0916] Terminal: Analyzes the input data and requests function suggestions from the server.
[0917] Step 3:
[0918] Server: Selects the most appropriate function or script based on the received data.
[0919] Step 4:
[0920] Server: Generates the selected functions and scripts as suggestions and sends them to the terminal.
[0921] Step 5:
[0922] Terminal: Automatically inserts the suggested function into the cell.
[0923] Step 6:
[0924] User: Review the proposed function and set the required data range and conditions.
[0925] Task organization and reminder function
[0926] Step 1:
[0927] User: Enter tasks into a task management app or scheduler.
[0928] Step 2:
[0929] Terminal: Sends entered task data to the server in real time.
[0930] Step 3:
[0931] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[0932] Step 4:
[0933] Server: Generates reminder notifications for identified tasks.
[0934] Step 5:
[0935] Server: Sends a reminder notification to the device.
[0936] Step 6:
[0937] Device: Displays reminder notifications to the user at the appropriate time.
[0938] Step 7:
[0939] User: Check the reminder notification and take action on the task.
[0940] (Example 1)
[0941] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0942] In modern work and daily life, there is a demand for efficiently managing and executing a wide range of tasks. In particular, there is a lack of tools to streamline tasks such as recording meeting content, clarifying the meaning of unfamiliar words, responding quickly to emails, analyzing complex data and proposing functions, and task reminders. Furthermore, users face the problem of having to spend a significant amount of time and effort performing these tasks individually.
[0943] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0944] In this invention, the server includes means for analyzing audio data and converting it into text, means for searching for the meaning and related information of detected words and phrases, means for generating a preliminary reply and displaying it as a draft on the terminal, means for analyzing data and generating appropriate functions and scripts, and means for analyzing task information and creating reminder notifications. This makes it possible to efficiently manage a wide range of tasks and significantly reduce the user's working time and effort.
[0945] "Audio data" refers to data that represents audio in a digital format.
[0946] A "server" is a computer system that processes and stores data over a network.
[0947] A "device" refers to a computer, smartphone, tablet, or other device used by a user.
[0948] A "user" is an individual who operates this system to perform work-related or daily life tasks.
[0949] A "microphone" is a device that collects sound and converts it into an electrical signal.
[0950] "Analysis" is the process of breaking down and interpreting data to reveal its contents.
[0951] "Text" is a collection of information expressed in written form.
[0952] "Real-time" refers to processing that occurs almost simultaneously without delay.
[0953] "Editing" is the process of modifying or changing text or other data.
[0954] "Saving" refers to the act of storing data in a storage device.
[0955] An "unknown word" is a word or phrase whose meaning is unknown to the user or system.
[0956] "Searching" is the process of finding specific information.
[0957] A "pop-up notification" is a notification window that temporarily appears on the screen.
[0958] A "first reply" is the initial draft of a response in communication such as email.
[0959] A "draft" is a preliminary version of a document created before it becomes a formal document.
[0960] A "function" is a mathematical or programming operation that takes a number or string as input and outputs a specific result.
[0961] A "script" is a series of commands or instructions that automatically execute a specific process.
[0962] A "task" is a specific activity or action in one's work or daily life.
[0963] A "reminder notification" is a notification that reminds you of the completion or deadline of a specific task.
[0964] A "generative AI model" is an artificial intelligence model that uses machine learning techniques to generate new data.
[0965] This invention relates to a resident generation AI system for users to efficiently manage and perform a wide range of tasks in their work and daily lives. This system provides real-time text conversion of voice data, notification of the meaning of unknown words, suggestions for automatic replies, suggestions for functions and scripts, and task reminder functions. Specific embodiments of this invention are described below.
[0966] System Configuration
[0967] Real-time transcription of meeting content
[0968] 1. Collection and transmission of audio data
[0969] Terminal: Collects audio in real time via the microphone during meetings and conversations. For example, use an audio capture library (e.g., PortAudio) on the terminal's operating system.
[0970] Terminal: Digitizes the collected audio data, performs encoding, and then sends it to the server using the HTTPS protocol.
[0971] 2. Analysis of audio data and text conversion
[0972] Server: Inputs the received audio data into a speech recognition API such as Google Cloud Speech-to-Text and converts it into text data.
[0973] 3. Temporary storage and real-time display of text data
[0974] Server: Temporarily stores the converted text data in a database (e.g., MySQL).
[0975] Terminal: Displays filtered text data in a dedicated window in real time. For example, it updates the page using JavaScript and WebSocket.
[0976] User: Review the displayed text and edit or save it as needed. The edited text will be saved to your local disk or cloud storage (e.g., Google Drive).
[0977] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server using Google Cloud Speech-to-Text. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[0978] Notification of the meaning of unfamiliar words
[0979] 1. Text input and analysis
[0980] User: Enter text into Excel or Google Docs.
[0981] Terminal: Analyzes the input text in real time and detects words that do not exist in the dictionary database (e.g., Oxford Dictionary API).
[0982] 2. Word meaning search and notification
[0983] Server: Uses the received word to call the Oxford Dictionaries API and retrieve its meaning and related information.
[0984] Device: Display search results as a pop-up notification. For example, using DOM manipulation and JavaScript.
[0985] User: Check the displayed explanation and obtain the necessary information.
[0986] Specific example: If a user tries to use the SUM function in Excel and doesn't understand its meaning, simply selecting the function name will cause the terminal to display a pop-up based on information from the Oxford Dictionaries. This allows the user to understand the meaning of the function without interrupting their work.
[0987] Suggestion for an automated reply
[0988] 1. Receiving and notifying emails
[0989] Terminal: Notifies the user that a new email has been received and displays its contents in the email client (e.g., Microsoft Outlook).
[0990] 2. Content analysis and generation of initial reply
[0991] Server: Analyzes the email content using a generation AI model (e.g., OpenAI GPT-4) and generates an initial reply.
[0992] 3. Draft display and confirmation
[0993] Terminal: Displays the generated initial reply as a draft.
[0994] User: Review the draft, make any necessary corrections, and send a reply.
[0995] Specific example: When a user receives a new email, the server analyzes its contents and generates an initial reply message such as, "Can you make any necessary adjustments?" The user can then review it, make any necessary corrections, and quickly send a reply.
[0996] Function suggestions for Excel and spreadsheets
[0997] 1. Data Input and Analysis
[0998] User: Enter data into Microsoft Excel or Google Sheets.
[0999] Terminal: Analyzes input data in real time. Using the Python pandas library is recommended.
[1000] 2. Function suggestions and input assistance
[1001] Server: Selects the most appropriate function based on the data content and generates a proposal using an AI model (e.g., OpenAI).
[1002] Terminal: Automatically inserts the suggested function into the cell.
[1003] User: Review the proposed function and set the data range and conditions.
[1004] Specific example: When a user aggregates sales data in a spreadsheet, a function like "=SUM(A2:A10)" is suggested and automatically inserted. The user can then easily check the data range and perform the calculation.
[1005] Task organization and reminder function
[1006] 1. Enter and submit task information
[1007] User: Enter task information into Microsoft To-Do or Google Calendar.
[1008] Terminal: Sends entered task information to the server in real time. Uses the HTTPS protocol.
[1009] 2. Task analysis and creation of reminder notifications
[1010] Server: Analyzes task information, extracts tasks with approaching deadlines, and creates reminder notifications.
[1011] 3. Displaying reminder notifications
[1012] Device: Displays reminder notifications to the user at the appropriate time.
[1013] User: Check reminder notifications and manage tasks.
[1014] Specific example: A user uses a task management app, and as the deadline for a task approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user receives this notification and can manage their tasks appropriately.
[1015] Examples of prompts for generative AI models
[1016] 1. "I want to collect the audio from the meeting and display it as text in real time. Please convert this audio to text."
[1017] 2. "I want to understand the meaning of functions used in Excel, so please explain the SUM function."
[1018] 3. "You have received a new email. Please generate an initial reply to this email immediately."
[1019] 4. "I'm entering sales data in Google Sheets, so please suggest an appropriate function."
[1020] 5. "I have set up reminder notifications in Google Calendar. The deadline for this task is approaching, so please create a reminder notification."
[1021] This invention's system allows users to efficiently carry out various tasks in their work and daily lives. It reduces the time and effort required for information retrieval, data entry, task management, etc., and provides a more efficient work environment.
[1022] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1023] Real-time transcription of meeting content
[1024] Step 1: Collect audio data
[1025] Terminal: Collects audio data via the microphone during meetings and conversations. This audio data is an analog signal and is sent from the microphone to the terminal's audio input port.
[1026] Input: Analog audio signal collected via microphone.
[1027] Output: An analog audio signal is sent to the terminal.
[1028] Step 2: Digitize and transmit audio data
[1029] Terminal: Performs an ADC (analog-to-digital converter) to convert analog audio signals into digital data. After conversion, encodes the digital audio data (e.g., MP3 or WAV) and sends it to the server using the HTTPS protocol.
[1030] Input: Analog audio signal.
[1031] Output: Digital audio data is sent to the server.
[1032] Step 3: Analyzing the audio data
[1033] Server: Inputs the received digital audio data into a speech recognition API (e.g., Google Cloud Speech-to-Text) for analysis. The API converts the audio data into text data.
[1034] Input: Digital audio data.
[1035] Output: Convert to text data.
[1036] Step 4: Save text data
[1037] Server: Temporarily stores the converted text data in a database (e.g., MySQL).
[1038] Input: Converted text data.
[1039] Output: Temporarily saved text data.
[1040] Step 5: Filter the text data.
[1041] Server: Extracts important keywords and phrases from stored text data and performs filtering to remove unnecessary noise. For example, natural language processing (NLP) techniques are used to extract content appropriate to the meeting context.
[1042] Input: Temporarily saved text data.
[1043] Output: Filtered text data.
[1044] Step 6: Real-time display of text
[1045] Terminal: Displays filtered text data in a dedicated window in real time. For example, it updates the page using JavaScript and WebSocket.
[1046] Input: Filtered text data.
[1047] Output: Text displayed in real time.
[1048] Step 7: Edit and save the text
[1049] User: Review the displayed text and edit or save it as needed. The edited text will be saved to your local disk or cloud storage (e.g., Google Drive).
[1050] Input: Text displayed in real time.
[1051] Output: Edited and saved text data.
[1052] Notification of the meaning of unfamiliar words
[1053] Step 1: Enter text
[1054] User: Enter text into Excel or Google Docs.
[1055] Input: Text entered by the user.
[1056] Output: Text is entered into the terminal.
[1057] Step 2: Detecting unknown words
[1058] Terminal: Analyzes the input text in real time and detects words that do not exist in the dictionary database (e.g., Oxford Dictionary API).
[1059] Input: The entered text.
[1060] Output: List of unknown words.
[1061] Step 3: Sending a word
[1062] Terminal: Sends unknown words to the server. For example, using an Ajax request.
[1063] Input: List of unknown words.
[1064] Output: Unknown word sent to the server.
[1065] Step 4: Semantic retrieval and analysis
[1066] Server: Uses a dictionary API to search for the meaning of the submitted word and related information, and analyzes the related information.
[1067] Input: An unknown word sent to the server.
[1068] Output: Meaning and related information.
[1069] Step 5: Notification of Meaning
[1070] Terminal: Displays the analyzed meaning and related information as a popup window. Specifically, it uses DOM manipulation and JavaScript.
[1071] Input: Meaning and related information.
[1072] Output: Meaning displayed in the pop-up notification.
[1073] Step 6: Confirm the contents
[1074] User: Check the displayed explanation and obtain the necessary information.
[1075] Input: Meaning of the pop-up notification.
[1076] Output: Information obtained by the user.
[1077] Suggestion for an automated reply
[1078] Step 1: Receiving and receiving emails
[1079] Terminal: Notifies the user that a new email has been received and displays its contents in the email client (e.g., Microsoft Outlook).
[1080] Input: Received email.
[1081] Output: Email notification displayed in the email client.
[1082] Step 2: Analyze the content of the email
[1083] Server: Sends the body of the received email to a generating AI model (e.g., OpenAI GPT-4) for analysis.
[1084] Input: The body of the received email.
[1085] Output: Analyzed email content.
[1086] Step 3: Generating the initial reply
[1087] Server: Based on the analysis results, it generates an initial reply message and applies its content to a template.
[1088] Input: Analyzed email content.
[1089] Output: Initial reply.
[1090] Step 4: Draft Display
[1091] Terminal: Displays the generated initial reply as a draft.
[1092] Input: Initial reply message.
[1093] Output: Reply displayed as a draft.
[1094] Step 5: Review and revise the draft
[1095] User: Review the draft and make any necessary corrections.
[1096] Input: Reply displayed as a draft.
[1097] Output: Revised reply text.
[1098] Step 6: Send a reply
[1099] User: Send the revised reply.
[1100] Input: Revised reply text.
[1101] Output: Sent reply email.
[1102] Function suggestions for Excel and spreadsheets
[1103] Step 1: Enter data
[1104] User: Enter data into Microsoft Excel or Google Sheets.
[1105] Input: Data entered by the user.
[1106] Output: Data entered into the terminal.
[1107] Step 2: Analysis of the month
[1108] Terminal: Analyzes input data in real time. Uses Python's pandas library, etc.
[1109] Input: The entered data.
[1110] Output: Analysis results.
[1111] Step 3: Propose a function
[1112] Server: Based on the data analysis results, it selects the optimal function and generates proposals using a generative AI model (e.g., OpenAI).
[1113] Input: Analysis results.
[1114] Output: Proposed function.
[1115] Step 4: Function display and input assistance
[1116] Terminal: Automatically inserts the suggested function into the cell.
[1117] Input: Proposed function.
[1118] Output: The function inserted into the cell.
[1119] Step 5: Verify and configure the function
[1120] User: Review the proposed function and set the required data range and conditions.
[1121] Input: The function inserted into the cell.
[1122] Output: The configured function.
[1123] Task organization and reminder function
[1124] Step 1: Enter task information
[1125] User: Enter task information into Microsoft To-Do or Google Calendar.
[1126] Input: Task information entered by the user.
[1127] Output: Task information entered into the terminal.
[1128] Step 2: Submit task data
[1129] Terminal: Sends entered task information to the server in real time. Uses the HTTPS protocol.
[1130] Input: Task information.
[1131] Output: Task information sent to the server.
[1132] Step 3: Task Analysis
[1133] Server: Analyzes task data and extracts tasks with approaching deadlines.
[1134] Input: Submitted task information.
[1135] Output: Extracted tasks with approaching deadlines.
[1136] Step 4: Create a reminder notification
[1137] Server: Creates reminder notifications for tasks with approaching deadlines.
[1138] Input: Extracted tasks with approaching deadlines.
[1139] Output: Reminder notification.
[1140] Step 5: Displaying reminder notifications
[1141] Terminal: Displays the created reminder notification to the user at the appropriate time.
[1142] Input: Reminder notification.
[1143] Output: Reminder notification displayed to the user.
[1144] Step 6: Task Management
[1145] User: Check reminder notifications and manage tasks.
[1146] Input: The displayed reminder notification.
[1147] Output: Managed tasks.
[1148] The above steps establish the processing flow of this system. This allows users to efficiently manage and perform various tasks and daily routines.
[1149] (Application Example 1)
[1150] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1151] In modern security operations, there is a demand for efficient information collection and management, as well as rapid response. However, manual reporting during patrols is time-consuming and laborious, and prone to information omissions and communication errors. Furthermore, daily tasks such as task reminders and initial email replies can add to the workload, reducing overall efficiency. In this context, there is a growing need for systems that convert speech to text in real time, automatically generate responses, and provide task reminders.
[1152] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1153] In this invention, the server includes means for collecting audio data through the microphone of a terminal used by the user, means for transmitting the collected audio data to the server, means for the server to analyze the audio data and convert it into text, means for generating an appropriate response sentence using a generation AI model based on the converted text, means for displaying the generated response sentence on the terminal in real time, means for transmitting task data registered by the user to the server, means for the server to analyze the task data and create a reminder notification for tasks with approaching deadlines, and means for displaying the reminder notification on the terminal. This enables real-time information gathering and response in security operations, thereby improving operational efficiency.
[1154] "User" refers to an individual or group that uses the system.
[1155] "Terminal" refers to a digital device used by a user (e.g., a smartphone or PC).
[1156] A "microphone" refers to a device that converts sound into electrical signals.
[1157] "Audio data" refers to the digital signal of sound collected through a microphone.
[1158] A "server" refers to a computer system that receives audio data, analyzes it, and converts it into text.
[1159] A "generative AI model" refers to artificial intelligence that generates appropriate responses or suggestions based on input text data.
[1160] "Text" refers to written information converted from audio data.
[1161] "Response sentence" refers to a dialogue-style reply sentence generated by a generative AI model.
[1162] "Task data" refers to information about tasks and schedules registered by the user.
[1163] A "reminder notification" refers to a notification that informs the user of tasks that are nearing their deadline.
[1164] This invention is a system for streamlining information gathering and task management in security operations. The system includes means for collecting voice data through the microphone of a terminal used by the user and transmitting the collected voice data to a server. The server analyzes the voice data, converts it to text, generates appropriate response sentences using a generative AI model, and displays them on the terminal in real time. Furthermore, the system transmits task data registered by the user to the server, which analyzes the task data, creates reminder notifications for tasks with approaching deadlines, and displays these on the terminal.
[1165] The server uses the Google Cloud Speech-to-Text API to convert speech data into text and generates response sentences based on the text data using a generative AI model (e.g., GPT-4). The terminal displays the transcribed meeting content and generated response sentences to the user in real time. This allows users to quickly and accurately obtain information and manage tasks during security work.
[1166] For example, if a night shift security staff member uses a mobile device to make a voice report during patrol, the voice is automatically converted to text and sent to the administrator in real time. Furthermore, reminder notifications are displayed based on a security checklist, helping to prevent missed tasks. This is expected to improve work efficiency and reduce errors.
[1167] The system of the present invention is specifically implemented through the following program processing. The program uses the Google Cloud Speech-to-Text API to convert speech data into text and generates response sentences using a generative AI model. Furthermore, a task reminder function is implemented using a scheduling library.
[1168] For example, the following prompt statements are possible:
[1169] "You are the designer of an application that helps streamline security services. It will transcribe voice data into text in real time and use AI to provide users with appropriate replies and task notifications. Please design an application with the following features:
[1170] Audio data collection and real-time text conversion
[1171] Generating appropriate automated replies from text data
[1172] "Recurring task reminder notifications"
[1173] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1174] Step 1:
[1175] The device collects audio data. The device collects the voice spoken by the user through the microphone in real time and saves it as data. In this step, the input is the user's voice, and the output is digitized audio data.
[1176] Step 2:
[1177] The terminal sends the collected audio data to the server. The collected audio data is uploaded to the server in real time. The input for this step is digitized audio data, and the output is audio data stored on the server.
[1178] Step 3:
[1179] The server analyzes the audio data and converts it to text. The server uses the Google Cloud Speech-to-Text API to convert the audio data to text data. The input for this step is audio data, and the output is text data. As a concrete example of how the audio analysis process works, an audio file is input to the API, and the result is returned as text.
[1180] Step 4:
[1181] The server generates an appropriate response sentence using a generative AI model based on the converted text. The generative AI model (e.g., GPT-4) is input with text data to generate a response sentence. The input for this step is text data, and the output is the generated response sentence. The generative AI model analyzes the text data to produce the most appropriate response or suggestion.
[1182] Step 5:
[1183] The server sends the generated response message to the terminal in real time. The generated response message is sent to the terminal and displayed to the user immediately. The input for this step is the response message, and the output is the response message displayed on the terminal.
[1184] Step 6:
[1185] The user registers task data on their device, and the device sends that data to the server. When a user registers a task using a task management app, that data is sent to the server in real time. In this step, the input is the task data entered by the user, and the output is the task data stored on the server.
[1186] Step 7:
[1187] The server analyzes task data and creates reminder notifications for tasks with approaching deadlines. It analyzes task data, extracts tasks with imminent deadlines, and generates reminder notifications. The input for this step is task data, and the output is the generated reminder notifications.
[1188] Step 8:
[1189] The server sends a reminder notification to the device, and the device displays it to the user. The server sends a reminder notification to the device and notifies the user. The input for this step is the reminder notification, and the output is the reminder notification displayed on the device.
[1190] The above processing steps enable real-time text conversion of audio data, generation of response statements, task management, and reminder notifications, thereby streamlining security operations.
[1191] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1192] This invention relates to a resident generative AI system for streamlining information retrieval, text conversion, task management, email correspondence, function input, and user emotion recognition and response in users' work and daily lives. This invention comprises the following components and their operation.
[1193] System Configuration
[1194] Real-time transcription of meeting content
[1195] Collection and transmission of audio data
[1196] Terminal: Collects audio during meetings in real time via the microphone.
[1197] Terminal: Digitizes the collected audio data and sends it to the server.
[1198] Audio data analysis and text conversion
[1199] Server: Inputs received audio data into a speech recognition model and converts it into text.
[1200] Server: Stores and filters the text-based data.
[1201] Server: Sends filtered text to the terminal.
[1202] Display text
[1203] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[1204] User: Review the displayed text and edit or save it as needed.
[1205] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[1206] Notification of the meaning of unfamiliar words
[1207] Text input and analysis
[1208] User: Enter text into Excel sheets or documents.
[1209] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[1210] Terminal: Sends detected words and phrases to the server.
[1211] Word meaning search
[1212] Server: Searches for the meaning of words and phrases, as well as related information.
[1213] Server: Create search results as a pop-up notification.
[1214] Meaning notification
[1215] Device: Displays a pop-up notification to the user.
[1216] User: Review the displayed explanation and access any further information you need.
[1217] Specific example: If a user doesn't understand the meaning of a particular function in Excel, selecting that function will display its meaning in a pop-up on the device. This allows the user to understand the meaning without interrupting their work.
[1218] Suggestion for an automated reply
[1219] Receiving and analyzing emails
[1220] Terminal: Notifies the user that a new email has been received.
[1221] Terminal: Sends the content of received emails to the server.
[1222] Reply generation and display
[1223] Server: Analyzes the content of the email and generates an initial reply.
[1224] Server: Sends the generated initial reply message to the terminal.
[1225] Terminal: Displays the initial reply as a draft.
[1226] User: Review the draft, make any necessary corrections, and then send a reply.
[1227] Specific example: When a user receives an email, the server analyzes its contents and generates an initial reply such as, "Could you please schedule a meeting on the following dates?" The user can then review and send the email, allowing for a quick response.
[1228] Function suggestions for Excel and spreadsheets
[1229] Text input and analysis
[1230] User: Enter data into Excel or a spreadsheet.
[1231] Terminal: Analyzes the input data and requests function suggestions from the server.
[1232] Function proposal generation and input assistance
[1233] Server: Selects the most appropriate function or script based on the received data.
[1234] Server: Generates the selected functions and scripts as suggestions and sends them to the terminal.
[1235] Terminal: Automatically inserts the suggested function into the cell.
[1236] User: Review the proposed function and set the required data range and conditions.
[1237] Specific example: When a user aggregates sales data, the "SUM function" is suggested, and the terminal automatically inserts the function. The user can perform the aggregation simply by specifying the data range.
[1238] Task organization and reminder function
[1239] Task data collection and transmission
[1240] User: Enter tasks into a task management app or scheduler.
[1241] Terminal: Sends entered task data to the server in real time.
[1242] Task analysis and reminders
[1243] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[1244] Server: Generates reminder notifications for identified tasks.
[1245] Server: Sends a reminder notification to the device.
[1246] Display reminder notifications
[1247] Device: Displays reminder notifications to the user at the appropriate time.
[1248] User: Check the reminder notification and take action on the task.
[1249] Specific example: When a user uses a task management app and the deadline approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user can then see the reminder and efficiently work on the task.
[1250] Additional configuration for the emotion engine
[1251] Recognition and response to emotions
[1252] Collection and analysis of voice and input data
[1253] Device: Collects voice and text input in real time.
[1254] Terminal: Sends collected data to the server.
[1255] Recognition of emotions
[1256] Server: Uses an emotion engine to analyze the user's emotional state from collected audio and text data.
[1257] Server: Classifies the user's emotional state based on the analysis results and determines appropriate feedback and responses.
[1258] Feedback and implementation of countermeasures
[1259] Server: Generates feedback and responses based on the user's emotional state and sends them to the terminal.
[1260] Terminal: Displays generated feedback and responses to the user.
[1261] Example 1: When a user receives an emotionally charged email, the emotion engine detects the user's stress level and suggests softening the tone of the reply. The device displays the suggestion, and the user can send it as is.
[1262] Example 2: If a user is working with a spreadsheet and the emotion engine detects the user's confusion or stress, the device will display appropriate guidance and support to help the user.
[1263] Specific example 3: In user task management, if the emotion engine detects user fatigue or anxiety, the server adjusts task priorities and sends a reminder notification to the device. The user can then check the notification and efficiently proceed with their tasks.
[1264] By combining the system of the present invention with an emotion engine, it becomes possible to recognize the user's emotional state and provide appropriate feedback and support. This allows users to work more comfortably and efficiently in their professional and daily lives.
[1265] The following describes the processing flow.
[1266] Additional configuration for the emotion engine
[1267] Recognition and response to emotions
[1268] Collection and analysis of voice and input data
[1269] Step 1:
[1270] Terminal: Collects audio from the user's voice through the microphone.
[1271] Step 2:
[1272] Terminal: Collects text data entered by the user.
[1273] Step 3:
[1274] Terminal: Sends collected audio and text data to the server.
[1275] Recognition of emotions
[1276] Step 1:
[1277] Server: Analyzes the collected audio data and determines the user's emotional state from the audio data.
[1278] Step 2:
[1279] Server: Analyzes collected text data and uses that data to analyze the user's emotional state.
[1280] Step 3:
[1281] Server: Integrates the analysis results of voice and text data to classify the user's overall emotional state.
[1282] Feedback and implementation of countermeasures
[1283] Step 1:
[1284] Server: Determines appropriate feedback and support based on emotional state.
[1285] Step 2:
[1286] Server: Generates the determined feedback and support details and sends them to the terminal.
[1287] Step 3:
[1288] Terminal: Displays generated feedback and support details to the user.
[1289] Specific example
[1290] Example 1: Emotional email responses
[1291] Step 1:
[1292] Terminal: Collects the content of emails received by the user and sends it to the server.
[1293] Step 2:
[1294] Server: Analyzes email content and uses an emotion engine to analyze the user's emotional state.
[1295] Step 3:
[1296] Server: Based on the analysis results, it generates a reply message to alleviate user stress.
[1297] Step 4:
[1298] Terminal: Displays the generated reply as a draft to the user.
[1299] Step 5:
[1300] User: Review the draft, make any necessary corrections, and then send a reply.
[1301] Example 2: Spreadsheet operation support
[1302] Step 1:
[1303] Terminal: Collects data entered by the user in a spreadsheet and sends it to the server.
[1304] Step 2:
[1305] Server: Uses an emotion engine to analyze the confusion and stress the user is experiencing.
[1306] Step 3:
[1307] Server: Generates appropriate guides and support content based on the analysis results.
[1308] Step 4:
[1309] Terminal: Displays generated guides and support information to the user.
[1310] Specific example 3: Adjusting task management and reminders
[1311] Step 1:
[1312] Terminal: Collects task data entered by the user into the task management app and sends it to the server.
[1313] Step 2:
[1314] Server: Uses an emotion engine to analyze user fatigue and anxiety.
[1315] Step 3:
[1316] Server: Based on the analysis results, adjust task priorities and reminder timings.
[1317] Step 4:
[1318] Server: Generates a coordinated reminder notification and sends it to the device.
[1319] Step 5:
[1320] Terminal: Displays the generated reminder notification to the user at the appropriate time.
[1321] Step 6:
[1322] User: Check the displayed reminder notification and take action on the task.
[1323] Thus, by combining an emotion engine, the system of the present invention can recognize the user's emotional state in real time and provide appropriate feedback and support, thereby improving efficiency in work and daily life.
[1324] (Example 2)
[1325] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1326] Modern information processing devices and various user systems require the efficient, real-time processing of a wide variety of tasks. However, centralized systems for quickly and accurately performing tasks such as converting audio data to text, searching for the meaning of unknown words and phrases, and automatically replying to emails are still not adequately developed. The lack of such systems can lead to decreased work efficiency and increased stress for users.
[1327] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing audio data and converting it into text using a speech recognition model, means for using the internet to search for the meaning of words and phrases and related information, and means for analyzing the content of emails and generating a primary reply using a generation AI model. This enables users to quickly record meeting content, instantly understand the meaning of unfamiliar words, and efficiently reply to emails.
[1328] "Audio data" refers to information that represents audio from meetings, conversations, etc., in a digital format.
[1329] "Digitalization" is the process of converting analog signals into digital signals.
[1330] A "server" is a computer that provides services and data to multiple clients over a network.
[1331] A "speech recognition model" is a machine learning model used to convert speech data into text.
[1332] "Text conversion" is the process of converting data such as audio or handwritten notes into text data.
[1333] "Filtering" is the process of applying specific conditions to data to extract only the necessary information.
[1334] An "information processing device" is an electronic computer that has the functions of inputting, processing, and outputting data.
[1335] An "unknown word or phrase" is a word or part of a sentence that the user cannot understand or find unclear.
[1336] A "pop-up notification" is a notification window that temporarily appears on the user's screen.
[1337] "Email" refers to messages that are sent and received electronically via the internet.
[1338] A "generative AI model" is a machine learning model used to generate text and data using artificial intelligence.
[1339] A "primary reply" is the initial response suggested in an email reply.
[1340] A "draft" is a preliminary version of a document or email before any revisions are made.
[1341] This invention relates to a resident generative AI system for streamlining information retrieval, text conversion, task management, email correspondence, function input, and user emotion recognition and response in users' work and daily lives. This invention comprises the following components and their operation.
[1342] Real-time transcription of meeting content
[1343] This system collects audio data in real time during meetings through the microphone of the user's information processing device, digitizes it, and sends it to a server. The server converts the received audio data into text using a speech recognition model (e.g., Google Cloud Speech-to-Text API). The transcribed data is stored in a database, where important keywords are filtered. The filtered text is sent to the information processing device in real time and displayed in a dedicated window.
[1344] Specific example:
[1345] When a user is conducting an online meeting, the information processing device collects audio data through its microphone and sends the digitized data to a server. The server uses a speech recognition model to transcribe the data into text, filters it, and then sends it back to the information processing device, allowing the meeting content to be displayed as text in real time.
[1346] Example of a prompt:
[1347] "Please explain, step by step, the process of the system that transcribes meeting content into text in real time."
[1348] Notification of the meaning of unfamiliar words
[1349] The system analyzes the text entered by the user into the information processing device in real time to detect unknown words and phrases. These detected words and phrases are sent to a server. The server searches for the meaning and related information of these words and phrases using the internet (e.g., Wikipedia API, Oxford Dictionaries API) and displays it on the information processing device as a pop-up notification.
[1350] Specific example:
[1351] When a user enters a specific function into an Excel sheet, it is detected and sent to the server. The server uses the corresponding API to look up the meaning of the function and displays it on the information processing device as a pop-up notification.
[1352] Example of a prompt:
[1353] "Please explain, step by step, the process of a system that automatically notifies you of the meaning of unknown words in Excel."
[1354] Suggestion for an automated reply
[1355] The server analyzes the content of the received email and generates a preliminary reply using a generative AI model (e.g., OpenAI GPT-3). The generated preliminary reply is displayed as a draft on the information processing device, and the user reviews it, makes corrections, and then sends it.
[1356] Specific example:
[1357] When a user receives an email, the server analyzes its contents and generates a preliminary reply such as, "Would you be able to schedule a meeting on the following dates?" The user can then review, revise, and send the reply, which is displayed as a draft.
[1358] Example of a prompt:
[1359] "Please explain the process of the system that generates automated reply messages, step by step."
[1360] Task organization and reminder function
[1361] When a user enters a task into a task management app or scheduler, that task data is sent to the server in real time. The server analyzes the task data, identifies unattended tasks with deadlines, and generates reminder notifications. These reminder notifications are sent to an information processing device and displayed to the user at the appropriate time.
[1362] Specific example:
[1363] When a user uses a task management app and the deadline approaches, the information processing device notifies them with a message such as, "Please complete this task by 3 PM tomorrow." The user can then use the reminder to efficiently complete the task.
[1364] Example of a prompt:
[1365] "Please describe the process of the system that sends task reminder notifications, step by step."
[1366] Recognition and response to emotions
[1367] This system incorporates an emotion engine that analyzes the user's emotional state using voice and text data collected from the user. This allows the system to classify the user's emotional state and determine appropriate feedback and responses. The generated feedback and responses are displayed on the information processing device.
[1368] Specific example:
[1369] When a user receives an emotionally charged email, the emotion engine detects the user's stress and suggests softening the tone of the reply. If the emotion engine detects confusion or stress while a user is working with a spreadsheet, the device displays appropriate guidance and support. In task management, if the emotion engine detects user fatigue or anxiety, the server adjusts task priorities and sends reminder notifications to the information processing device.
[1370] Example of a prompt:
[1371] "Please describe, step by step, how the system recognizes and appropriately responds to user emotions."
[1372] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1373] Real-time transcription of meeting content
[1374] Step 1:
[1375] Voice collection
[1376] Device: Collects audio during meetings in real time using a microphone.
[1377] Input: Analog audio from a meeting.
[1378] Output: Analog audio data.
[1379] Specific operation: By pressing the "Start Recording" button, the device activates the microphone and begins collecting audio data.
[1380] Step 2:
[1381] Audio digitization
[1382] Terminal: Digitizes the collected analog audio data.
[1383] Input: Analog audio data.
[1384] Output: Digital audio data.
[1385] Specific operation: Digitize the audio at a sampling rate of 44.1kHz and convert it to PCM format.
[1386] Step 3:
[1387] Sending audio data
[1388] Terminal: Sends digitized audio data to the server.
[1389] Input: Digital audio data.
[1390] Output: Transmission of digital audio data to the server.
[1391] Specific operation: Digitized audio data is sent to the server using a real-time streaming protocol (e.g., WebSocket).
[1392] Step 4:
[1393] Receiving audio data
[1394] Server: The server receives audio data in real time.
[1395] Input: Digital audio data transmitted from the device.
[1396] Output: Audio data stored in the buffer.
[1397] Specific operation: The server continuously receives frames of audio data and stores them in a fixed buffer.
[1398] Step 5:
[1399] Speech recognition and text conversion
[1400] Server: Converts received audio data into text using a speech recognition model.
[1401] Input: Audio data stored in the buffer.
[1402] Output: Text data.
[1403] Specific operation: When the audio data buffer reaches a certain amount, the data is sent to the speech recognition API as a request, and the text result is received.
[1404] Step 6:
[1405] Text filtering and saving
[1406] Server: Stores the digitized data in a database and filters for important keywords.
[1407] Input: Text data.
[1408] Output: Filtered text data.
[1409] Specific operation: When saving text to the database, an SQL query is used to set the "importance" field, and if a specific keyword is included, it is set to "high," etc.
[1410] Step 7:
[1411] Sending filtered text
[1412] Server: Transmits text data to the information processing device in real time.
[1413] Input: Filtered text data.
[1414] Output: Sending text data to an information processing device.
[1415] Specific operation: The server sends filtered text data to the terminal via WebSocket.
[1416] Step 8:
[1417] Display text
[1418] Terminal: Displays text data in real time in a dedicated window.
[1419] Input: Filtered text data.
[1420] Output: The text displayed to the user.
[1421] Specific operation: Uses Javascript and other front-end technologies to manipulate the DOM within the window and render text.
[1422] Step 9:
[1423] Edit and save text
[1424] User: Review the displayed text, edit it as needed, and save it.
[1425] Input: The displayed text.
[1426] Output: Saved text data.
[1427] Specific operation: A text editor is provided on the editing screen, and the data is saved to local storage or cloud storage by pressing the "Save" button.
[1428] Notification of the meaning of unfamiliar words
[1429] Step 1:
[1430] Text input
[1431] User: Enter text into Excel or other documents.
[1432] Input: Text entered by the user.
[1433] Output: Input text data.
[1434] Specific action: The user enters formulas and explanations into an Excel sheet.
[1435] Step 2:
[1436] Real-time analysis
[1437] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[1438] Input: The entered text data.
[1439] Output: Unknown words or phrases detected.
[1440] Specific operation: Every time text is entered, it is analyzed using NLP (e.g., spaCy).
[1441] Step 3:
[1442] Sending to the server
[1443] Terminal: Sends detected words and phrases to the server.
[1444] Input: Unknown words or phrases detected.
[1445] Output: Data sent to the server.
[1446] Specific operation: Encode the detected data into JSON format and send it to the server via an HTTP request.
[1447] Step 4:
[1448] Search for meaning
[1449] Server: The server searches for the meaning of words and phrases using the internet.
[1450] Input: Unknown words or phrases detected.
[1451] Output: Search results data.
[1452] Specific operation: Send requests to the Wikipedia API and other dictionary APIs to retrieve meanings and related information.
[1453] Step 5:
[1454] Pop-up notification generation
[1455] Server: Generates a popup notification based on the search results.
[1456] Input: Search results data.
[1457] Output: Popup notification data.
[1458] Specific operation: Format the retrieved search results and generate a popup notification using HTML and CSS.
[1459] Step 6:
[1460] Send to device
[1461] Server: Sends a pop-up notification to the device.
[1462] Input: Pop-up notification data.
[1463] Output: Data sent to the terminal.
[1464] Specific action: The created popup notification is sent to the device as an HTTP response in JSON format.
[1465] Step 7:
[1466] Pop-up display
[1467] Device: Displays a pop-up notification on the user's screen.
[1468] Input: Pop-up notification data from the server.
[1469] Output: A pop-up notification displayed on the screen.
[1470] Specific operation: Use JavaScript to display a popup at the appropriate location and timing.
[1471] Suggestion for an automated reply
[1472] Step 1:
[1473] Email notification
[1474] Terminal: Notifies the user that a new email has been received.
[1475] Input: Received email.
[1476] Output: Notification of incoming email.
[1477] Specific action: Notify the user via a pop-up notification that a new email has been received.
[1478] Step 2:
[1479] Sending the email content
[1480] Terminal: Sends the content of received emails to the server.
[1481] Input: Received email data.
[1482] Output: Data sent to the server.
[1483] Specific operation: Convert the email content to JSON format and send it to the server via an HTTP request.
[1484] Step 3:
[1485] Generating a reply
[1486] Server: The server analyzes the email content and generates an initial reply using a generation AI model.
[1487] Input: Email content data.
[1488] Output: Initial reply.
[1489] Specific operation: Analyze the content of the email and generate a reply using a generative AI model such as OpenAI GPT-3.
[1490] Step 4:
[1491] Sending a reply
[1492] Server: Sends the generated initial reply to the terminal as a draft.
[1493] Input: Initial reply.
[1494] Output: Data sent to the terminal.
[1495] Specific operation: Convert the reply text to JSON format and send it to the terminal as an HTTP response.
[1496] Step 5:
[1497] Displaying the reply
[1498] Terminal: Displays the initial reply as a draft.
[1499] Input: Initial response data from the server.
[1500] Output: Draft below the user.
[1501] Specific operation: Use JavaScript to display a draft of the reply message in the browser.
[1502] Step 6:
[1503] User verification and correction
[1504] User: Review the draft and make any necessary corrections.
[1505] Input: Draft of the initial reply.
[1506] Output: Revised reply.
[1507] Specific operation: The user makes modifications through a text editor in the browser.
[1508] Step 7:
[1509] Corrected and sent
[1510] User: Send the revised reply.
[1511] Input: Revised reply.
[1512] Output: Sent email.
[1513] Specific action: Press the "Send" button to send the revised reply to the recipient.
[1514] The above steps demonstrate how each function is specifically implemented. Clearly defining the specific actions and data flows performed at each processing step makes it easier to understand the overall operation of the system.
[1515] (Application Example 2)
[1516] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1517] In modern brick-and-mortar stores, employees are expected to respond to customers quickly and accurately. However, employees may be unable to provide appropriate service due to lack of experience or information. Furthermore, understanding and responding appropriately to customer emotions is also difficult. A system is needed that allows employees to receive information in real time and provide the best possible response based on customer emotions.
[1518] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing sound and text information and recognizing the emotional state of the user, means for presenting feedback and responses based on the emotional state, means for optimizing dialogue with the customer using the emotional analysis results, and means for providing support information to a visual display device for quickly and appropriately answering customer questions. This enables employees in physical stores to respond to customers quickly and appropriately.
[1519] An "information processing device" is a device used to process, store, and transmit data, and includes smartphones, tablets, and computers.
[1520] "Sound" is a wave produced by vibrations in the air, and is an audio signal collected through an input device such as a microphone.
[1521] A "network" is a communication system that allows multiple information processing devices to send and receive data from one another.
[1522] "Textual information" refers to data obtained by analyzing audio or handwritten data and converting it into text format.
[1523] "Real-time" means that data is processed and displayed almost simultaneously, with virtually no delay.
[1524] "Emotional state" refers to the emotional state of a user as analyzed using an emotion engine, and includes emotions such as joy, anger, sadness, and surprise.
[1525] "Feedback" refers to information and advice provided based on a user's behavior and circumstances.
[1526] "Response" refers to appropriate actions or measures taken in response to a specific situation or condition.
[1527] A "visual display device" is a device used to visually display data, and includes smart glasses, displays, and head-mounted displays.
[1528] A "customer" is someone who intends to purchase or use a product or service.
[1529] "Support information" refers to information and data that are provided to help users perform specific tasks.
[1530] "Emotional analysis results" refer to data obtained after evaluating and classifying emotional states.
[1531] This invention provides an information processing system for improving the customer service capabilities of employees in physical stores. The embodiments for carrying out this invention are described in detail below.
[1532] The entire system consists of a visual display device used by the user (such as smart glasses), a microphone for collecting audio, a server for analyzing the data, and a network to connect them.
[1533] 1. Speech recognition and transcription
[1534] First, the user wears smart glasses, and questions from customers in the store are collected via a microphone. The visual display device acquires the audio and transmits it to a server over the network. The server analyzes the audio data and converts it into text using a speech recognition model.
[1535] 2. Recognition of emotional states
[1536] The server analyzes customer facial expression data acquired through visual display devices and cameras, and uses an emotion engine to recognize the customer's emotional state. Emotional states include feelings such as joy, anger, sadness, and surprise. This allows the user to understand the customer's feelings.
[1537] 3. Feedback and proposed response
[1538] The server provides appropriate feedback and responses based on the recognized emotional state of the customer. For example, if the customer is angry, it suggests a response in a calm tone. If the customer is confused, it provides detailed explanations and guidance.
[1539] 4. Provision of support information
[1540] In response to customer inquiries, the server generates appropriate answers, which are displayed in real time on the smart glasses. This allows users to instantly answer customer questions. Other relevant information and product details are also presented as needed.
[1541] These features enable employees to respond to customers quickly and accurately, leading to improved customer satisfaction.
[1542] Examples
[1543] One day, if a customer in a physical store asks, "How much does this item cost?", the microphone in the smart glasses captures the voice and transmits it to a server via the network. The server converts the voice into text and also recognizes the customer's emotions from their facial expressions. Based on the emotional state, the server provides a calm response, such as, "This item costs 299 yen." This information is displayed in real time on the smart glasses' screen, allowing employees to respond to customers quickly.
[1544] Example of a prompt
[1545] "A customer has asked a question about a product. Based on that question and the subsequent conversation, please have the AI generate an appropriate answer. Also, please suggest responses that take into account the customer's emotional state."
[1546] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1547] Step 1:
[1548] The user wears smart glasses and receives customer questions via voice. A microphone built into the smart glasses collects the voice data. This voice data is input into the user's device (smart glasses).
[1549] Step 2:
[1550] The terminal digitizes the collected audio data and transmits it to a server via the network. Here, the data format is converted, and the audio data is input to the server.
[1551] Step 3:
[1552] The server inputs the received audio data into a speech recognition model (e.g., Google Speech-to-Text API) and converts it into text information. In this process, the audio data is processed into text data, and the result is output to the server.
[1553] Step 4:
[1554] The server analyzes the converted text information and extracts specific keywords (e.g., product name, price). Furthermore, it searches for supporting information (such as price information from a database) to generate the most suitable response based on that information. These search results are then output to the server.
[1555] Step 5:
[1556] The server uses the smart glasses' camera to acquire customer facial expression data and analyzes it in real time. Using an emotion recognition engine such as DeepFace, it classifies and recognizes the customer's emotional state (e.g., anger, joy). As a result, the emotional data is output to the server.
[1557] Step 6:
[1558] The server generates an action plan based on the recognized emotional state and the content of the question. For example, if the customer is angry, it will suggest a polite way to handle the situation. This information is then sent from the server to the terminal as feedback.
[1559] Step 7:
[1560] The device (smart glasses) displays feedback information (responses and suggested actions) sent from the server in real time. The user then uses this display to respond to or suggest to the customer.
[1561] Step 8:
[1562] When a user provides an appropriate response or suggestion to a customer, the terminal additionally records that information and sends it to the server as needed to improve the relevant data. This allows the system to learn and further improve its response in the future.
[1563] This processing flow is expected to significantly improve the customer service capabilities of employees in physical stores, leading to increased customer satisfaction.
[1564] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1565] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1566] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1567] [Third Embodiment]
[1568] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1569] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1570] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1571] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1572] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1573] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1574] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1575] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1576] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1577] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1578] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1579] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1580] This invention relates to a resident generation AI system for streamlining information retrieval, text conversion, task management, email correspondence, function input, and other tasks in users' work and daily lives. This invention comprises the following components and their operation.
[1581] System Configuration
[1582] Real-time transcription of meeting content
[1583] Collection and transmission of audio data
[1584] Terminal: Collects audio in real time via microphone during meetings and conversations.
[1585] Terminal: Digitizes the collected audio data and sends it to the server.
[1586] Audio data analysis and text conversion
[1587] Server: Inputs the received audio data into the speech recognition model and converts it into text.
[1588] Server: Temporarily stores the text-based data and provides the user with appropriately filtered content.
[1589] Display text
[1590] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[1591] User: Review the displayed text and edit or save it as needed.
[1592] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[1593] Notification of the meaning of unfamiliar words
[1594] Text input and analysis
[1595] User: Enter text into Excel sheets or documents.
[1596] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[1597] Word meaning search
[1598] Server: Searches for the meaning of a word or phrase selected by the user from related databases.
[1599] Server: Analyzes relevant information and extracts appropriate content.
[1600] Meaning notification
[1601] Device: Display search results as a pop-up notification.
[1602] User: Check the displayed explanation and quickly obtain the necessary information.
[1603] Specific example: If a user doesn't understand the meaning of a particular function in Excel, selecting that function will display its meaning in a pop-up on the device. This allows the user to understand the meaning without interrupting their work.
[1604] Suggestion for an automated reply
[1605] Receiving and analyzing emails
[1606] Terminal: Notifies the user that a new email has been received and displays its contents.
[1607] Server: Analyzes the content of the email and generates an appropriate initial reply.
[1608] Reply generation and display
[1609] Server: Generates an initial reply and applies it to the template.
[1610] Terminal: Displays the generated initial reply as a draft.
[1611] User: Review the draft, make any necessary corrections, and then send a reply.
[1612] Specific example: When a user receives an email, the server analyzes its contents and generates an initial reply such as, "Could you please schedule a meeting on the following dates?" The user can then review and send the email, allowing for a quick response.
[1613] Function suggestions for Excel and spreadsheets
[1614] Text input and analysis
[1615] User: Enter data into Excel or a spreadsheet.
[1616] Terminal: Analyzes the input data and suggests appropriate functions or scripts.
[1617] Function proposal generation and input assistance
[1618] Server: Based on the input, it selects the most appropriate function or script and generates it as a suggestion.
[1619] Terminal: Automatically inserts the suggested function into the cell.
[1620] User: Review the proposed function and set the required data range and conditions.
[1621] Specific example: When a user aggregates sales data, the "SUM function" is suggested, and the terminal automatically inserts the function. The user can perform the aggregation simply by specifying the data range.
[1622] Task organization and reminder function
[1623] Task data collection and transmission
[1624] User: Enter tasks into a task management app or scheduler.
[1625] Terminal: Sends entered task data to the server in real time.
[1626] Task analysis and reminders
[1627] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[1628] Server: Creates reminder notifications for tasks with approaching deadlines.
[1629] Display reminder notifications
[1630] Device: Displays reminder notifications to the user at the appropriate time.
[1631] User: Check the reminder notification and take action on the task.
[1632] Specific example: When a user uses a task management app and the deadline approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user can then see the reminder and efficiently work on the task.
[1633] The system of the present invention allows users to efficiently carry out various tasks in their work and daily lives, significantly reducing the time and effort required for information retrieval, data entry, task management, and other related activities.
[1634] The following describes the processing flow.
[1635] Real-time transcription of meeting content
[1636] Step 1:
[1637] Terminal: Collects audio during meetings in real time via the microphone.
[1638] Step 2:
[1639] Terminal: Digitizes the collected audio data and sends it to the server.
[1640] Step 3:
[1641] Server: Inputs received audio data into a speech recognition model and converts it into text.
[1642] Step 4:
[1643] Server: Stores and filters the text-based data.
[1644] Step 5:
[1645] Server: Sends filtered text to the terminal.
[1646] Step 6:
[1647] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[1648] Step 7:
[1649] User: Review the displayed text and edit or save it as needed.
[1650] Notification of the meaning of unfamiliar words
[1651] Step 1:
[1652] User: Enter text into Excel sheets or documents.
[1653] Step 2:
[1654] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[1655] Step 3:
[1656] Terminal: Sends detected words and phrases to the server.
[1657] Step 4:
[1658] Server: Searches for the meaning of words and phrases, as well as related information.
[1659] Step 5:
[1660] Server: Create search results as a pop-up notification.
[1661] Step 6:
[1662] Device: Displays a pop-up notification to the user.
[1663] Step 7:
[1664] User: Review the displayed explanation and access any further information you need.
[1665] Suggestion for an automated reply
[1666] Step 1:
[1667] Terminal: Notifies the user that a new email has been received.
[1668] Step 2:
[1669] Terminal: Sends the content of received emails to the server.
[1670] Step 3:
[1671] Server: Analyzes the content of the email and generates an initial reply.
[1672] Step 4:
[1673] Server: Sends the generated initial reply message to the terminal.
[1674] Step 5:
[1675] Terminal: Displays the initial reply as a draft.
[1676] Step 6:
[1677] User: Review the draft, make any necessary corrections, and then submit.
[1678] Function suggestions for Excel and spreadsheets
[1679] Step 1:
[1680] User: Enter data into Excel or a spreadsheet.
[1681] Step 2:
[1682] Terminal: Analyzes the input data and requests function suggestions from the server.
[1683] Step 3:
[1684] Server: Selects the most appropriate function or script based on the received data.
[1685] Step 4:
[1686] Server: Generates the selected functions and scripts as suggestions and sends them to the terminal.
[1687] Step 5:
[1688] Terminal: Automatically inserts the suggested function into the cell.
[1689] Step 6:
[1690] User: Review the proposed function and set the required data range and conditions.
[1691] Task organization and reminder function
[1692] Step 1:
[1693] User: Enter tasks into a task management app or scheduler.
[1694] Step 2:
[1695] Terminal: Sends entered task data to the server in real time.
[1696] Step 3:
[1697] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[1698] Step 4:
[1699] Server: Generates reminder notifications for identified tasks.
[1700] Step 5:
[1701] Server: Sends a reminder notification to the device.
[1702] Step 6:
[1703] Device: Displays reminder notifications to the user at the appropriate time.
[1704] Step 7:
[1705] User: Check the reminder notification and take action on the task.
[1706] (Example 1)
[1707] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1708] In modern work and daily life, there is a demand for efficiently managing and executing a wide range of tasks. In particular, there is a lack of tools to streamline tasks such as recording meeting content, clarifying the meaning of unfamiliar words, responding quickly to emails, analyzing complex data and proposing functions, and task reminders. Furthermore, users face the problem of having to spend a significant amount of time and effort performing these tasks individually.
[1709] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1710] In this invention, the server includes means for analyzing audio data and converting it into text, means for searching for the meaning and related information of detected words and phrases, means for generating a preliminary reply and displaying it as a draft on the terminal, means for analyzing data and generating appropriate functions and scripts, and means for analyzing task information and creating reminder notifications. This makes it possible to efficiently manage a wide range of tasks and significantly reduce the user's working time and effort.
[1711] "Audio data" refers to data that represents audio in a digital format.
[1712] A "server" is a computer system that processes and stores data over a network.
[1713] A "device" refers to a computer, smartphone, tablet, or other device used by a user.
[1714] A "user" is an individual who operates this system to perform work-related or daily life tasks.
[1715] A "microphone" is a device that collects sound and converts it into an electrical signal.
[1716] "Analysis" is the process of breaking down and interpreting data to reveal its contents.
[1717] "Text" is a collection of information expressed in written form.
[1718] "Real-time" refers to processing that occurs almost simultaneously without delay.
[1719] "Editing" is the process of modifying or changing text or other data.
[1720] "Saving" refers to the act of storing data in a storage device.
[1721] An "unknown word" is a word or phrase whose meaning is unknown to the user or system.
[1722] "Searching" is the process of finding specific information.
[1723] A "pop-up notification" is a notification window that temporarily appears on the screen.
[1724] A "first reply" is the initial draft of a response in communication such as email.
[1725] A "draft" is a preliminary version of a document created before it becomes a formal document.
[1726] A "function" is a mathematical or programming operation that takes a number or string as input and outputs a specific result.
[1727] A "script" is a series of commands or instructions that automatically execute a specific process.
[1728] A "task" is a specific activity or action in one's work or daily life.
[1729] A "reminder notification" is a notification that reminds you of the completion or deadline of a specific task.
[1730] A "generative AI model" is an artificial intelligence model that uses machine learning techniques to generate new data.
[1731] This invention relates to a resident generation AI system for users to efficiently manage and perform a wide range of tasks in their work and daily lives. This system provides real-time text conversion of voice data, notification of the meaning of unknown words, suggestions for automatic replies, suggestions for functions and scripts, and task reminder functions. Specific embodiments of this invention are described below.
[1732] System Configuration
[1733] Real-time transcription of meeting content
[1734] 1. Collection and transmission of audio data
[1735] Terminal: Collects audio in real time via the microphone during meetings and conversations. For example, use an audio capture library (e.g., PortAudio) on the terminal's operating system.
[1736] Terminal: Digitizes the collected audio data, performs encoding, and then sends it to the server using the HTTPS protocol.
[1737] 2. Analysis of audio data and text conversion
[1738] Server: Inputs the received audio data into a speech recognition API such as Google Cloud Speech-to-Text and converts it into text data.
[1739] 3. Temporary storage and real-time display of text data
[1740] Server: Temporarily stores the converted text data in a database (e.g., MySQL).
[1741] Terminal: Displays filtered text data in a dedicated window in real time. For example, it updates the page using JavaScript and WebSocket.
[1742] User: Review the displayed text and edit or save it as needed. The edited text will be saved to your local disk or cloud storage (e.g., Google Drive).
[1743] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server using Google Cloud Speech-to-Text. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[1744] Notification of the meaning of unfamiliar words
[1745] 1. Text input and analysis
[1746] User: Enter text into Excel or Google Docs.
[1747] Terminal: Analyzes the input text in real time and detects words that do not exist in the dictionary database (e.g., Oxford Dictionary API).
[1748] 2. Word meaning search and notification
[1749] Server: Uses the received word to call the Oxford Dictionaries API and retrieve its meaning and related information.
[1750] Device: Display search results as a pop-up notification. For example, using DOM manipulation and JavaScript.
[1751] User: Check the displayed explanation and obtain the necessary information.
[1752] Specific example: If a user tries to use the SUM function in Excel and doesn't understand its meaning, simply selecting the function name will cause the terminal to display a pop-up based on information from the Oxford Dictionaries. This allows the user to understand the meaning of the function without interrupting their work.
[1753] Suggestion for an automated reply
[1754] 1. Receiving and notifying emails
[1755] Terminal: Notifies the user that a new email has been received and displays its contents in the email client (e.g., Microsoft Outlook).
[1756] 2. Content analysis and generation of initial reply
[1757] Server: Analyzes the email content using a generation AI model (e.g., OpenAI GPT-4) and generates an initial reply.
[1758] 3. Draft display and confirmation
[1759] Terminal: Displays the generated initial reply as a draft.
[1760] User: Review the draft, make any necessary corrections, and send a reply.
[1761] Specific example: When a user receives a new email, the server analyzes its contents and generates an initial reply message such as, "Can you make any necessary adjustments?" The user can then review it, make any necessary corrections, and quickly send a reply.
[1762] Function suggestions for Excel and spreadsheets
[1763] 1. Data Input and Analysis
[1764] User: Enter data into Microsoft Excel or Google Sheets.
[1765] Terminal: Analyzes input data in real time. Using the Python pandas library is recommended.
[1766] 2. Function suggestions and input assistance
[1767] Server: Selects the most appropriate function based on the data content and generates a proposal using an AI model (e.g., OpenAI).
[1768] Terminal: Automatically inserts the suggested function into the cell.
[1769] User: Review the proposed function and set the data range and conditions.
[1770] Specific example: When a user aggregates sales data in a spreadsheet, a function like "=SUM(A2:A10)" is suggested and automatically inserted. The user can then easily check the data range and perform the calculation.
[1771] Task organization and reminder function
[1772] 1. Enter and submit task information
[1773] User: Enter task information into Microsoft To-Do or Google Calendar.
[1774] Terminal: Sends entered task information to the server in real time. Uses the HTTPS protocol.
[1775] 2. Task analysis and creation of reminder notifications
[1776] Server: Analyzes task information, extracts tasks with approaching deadlines, and creates reminder notifications.
[1777] 3. Displaying reminder notifications
[1778] Device: Displays reminder notifications to the user at the appropriate time.
[1779] User: Check reminder notifications and manage tasks.
[1780] Specific example: A user uses a task management app, and as the deadline for a task approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user receives this notification and can manage their tasks appropriately.
[1781] Examples of prompts for generative AI models
[1782] 1. "I want to collect the audio from the meeting and display it as text in real time. Please convert this audio to text."
[1783] 2. "I want to understand the meaning of functions used in Excel, so please explain the SUM function."
[1784] 3. "You have received a new email. Please generate an initial reply to this email immediately."
[1785] 4. "I'm entering sales data in Google Sheets, so please suggest an appropriate function."
[1786] 5. "I have set up reminder notifications in Google Calendar. The deadline for this task is approaching, so please create a reminder notification."
[1787] This invention's system allows users to efficiently carry out various tasks in their work and daily lives. It reduces the time and effort required for information retrieval, data entry, task management, etc., and provides a more efficient work environment.
[1788] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1789] Real-time transcription of meeting content
[1790] Step 1: Collect audio data
[1791] Terminal: Collects audio data via the microphone during meetings and conversations. This audio data is an analog signal and is sent from the microphone to the terminal's audio input port.
[1792] Input: Analog audio signal collected via microphone.
[1793] Output: An analog audio signal is sent to the terminal.
[1794] Step 2: Digitize and transmit audio data
[1795] Terminal: Performs an ADC (analog-to-digital converter) to convert analog audio signals into digital data. After conversion, encodes the digital audio data (e.g., MP3 or WAV) and sends it to the server using the HTTPS protocol.
[1796] Input: Analog audio signal.
[1797] Output: Digital audio data is sent to the server.
[1798] Step 3: Analyzing the audio data
[1799] Server: Inputs the received digital audio data into a speech recognition API (e.g., Google Cloud Speech-to-Text) for analysis. The API converts the audio data into text data.
[1800] Input: Digital audio data.
[1801] Output: Convert to text data.
[1802] Step 4: Save text data
[1803] Server: Temporarily stores the converted text data in a database (e.g., MySQL).
[1804] Input: Converted text data.
[1805] Output: Temporarily saved text data.
[1806] Step 5: Filter the text data.
[1807] Server: Extracts important keywords and phrases from stored text data and performs filtering to remove unnecessary noise. For example, natural language processing (NLP) techniques are used to extract content appropriate to the meeting context.
[1808] Input: Temporarily saved text data.
[1809] Output: Filtered text data.
[1810] Step 6: Real-time display of text
[1811] Terminal: Displays filtered text data in a dedicated window in real time. For example, it updates the page using JavaScript and WebSocket.
[1812] Input: Filtered text data.
[1813] Output: Text displayed in real time.
[1814] Step 7: Edit and save the text
[1815] User: Review the displayed text and edit or save it as needed. The edited text will be saved to your local disk or cloud storage (e.g., Google Drive).
[1816] Input: Text displayed in real time.
[1817] Output: Edited and saved text data.
[1818] Notification of the meaning of unfamiliar words
[1819] Step 1: Enter text
[1820] User: Enter text into Excel or Google Docs.
[1821] Input: Text entered by the user.
[1822] Output: Text is entered into the terminal.
[1823] Step 2: Detecting unknown words
[1824] Terminal: Analyzes the input text in real time and detects words that do not exist in the dictionary database (e.g., Oxford Dictionary API).
[1825] Input: The entered text.
[1826] Output: List of unknown words.
[1827] Step 3: Sending a word
[1828] Terminal: Sends unknown words to the server. For example, using an Ajax request.
[1829] Input: List of unknown words.
[1830] Output: Unknown word sent to the server.
[1831] Step 4: Semantic retrieval and analysis
[1832] Server: Uses a dictionary API to search for the meaning of the submitted word and related information, and analyzes the related information.
[1833] Input: An unknown word sent to the server.
[1834] Output: Meaning and related information.
[1835] Step 5: Notification of Meaning
[1836] Terminal: Displays the analyzed meaning and related information as a popup window. Specifically, it uses DOM manipulation and JavaScript.
[1837] Input: Meaning and related information.
[1838] Output: Meaning displayed in the pop-up notification.
[1839] Step 6: Confirm the contents
[1840] User: Check the displayed explanation and obtain the necessary information.
[1841] Input: Meaning of the pop-up notification.
[1842] Output: Information obtained by the user.
[1843] Suggestion for an automated reply
[1844] Step 1: Receiving and receiving emails
[1845] Terminal: Notifies the user that a new email has been received and displays its contents in the email client (e.g., Microsoft Outlook).
[1846] Input: Received email.
[1847] Output: Email notification displayed in the email client.
[1848] Step 2: Analyze the content of the email
[1849] Server: Sends the body of the received email to a generating AI model (e.g., OpenAI GPT-4) for analysis.
[1850] Input: The body of the received email.
[1851] Output: Analyzed email content.
[1852] Step 3: Generating the initial reply
[1853] Server: Based on the analysis results, it generates an initial reply message and applies its content to a template.
[1854] Input: Analyzed email content.
[1855] Output: Initial reply.
[1856] Step 4: Draft Display
[1857] Terminal: Displays the generated initial reply as a draft.
[1858] Input: Initial reply message.
[1859] Output: Reply displayed as a draft.
[1860] Step 5: Review and revise the draft
[1861] User: Review the draft and make any necessary corrections.
[1862] Input: Reply displayed as a draft.
[1863] Output: Revised reply text.
[1864] Step 6: Send a reply
[1865] User: Send the revised reply.
[1866] Input: Revised reply text.
[1867] Output: Sent reply email.
[1868] Function suggestions for Excel and spreadsheets
[1869] Step 1: Enter data
[1870] User: Enter data into Microsoft Excel or Google Sheets.
[1871] Input: Data entered by the user.
[1872] Output: Data entered into the terminal.
[1873] Step 2: Analysis of the month
[1874] Terminal: Analyzes input data in real time. Uses Python's pandas library, etc.
[1875] Input: The entered data.
[1876] Output: Analysis results.
[1877] Step 3: Propose a function
[1878] Server: Based on the data analysis results, it selects the optimal function and generates proposals using a generative AI model (e.g., OpenAI).
[1879] Input: Analysis results.
[1880] Output: Proposed function.
[1881] Step 4: Function display and input assistance
[1882] Terminal: Automatically inserts the suggested function into the cell.
[1883] Input: Proposed function.
[1884] Output: The function inserted into the cell.
[1885] Step 5: Verify and configure the function
[1886] User: Review the proposed function and set the required data range and conditions.
[1887] Input: The function inserted into the cell.
[1888] Output: The configured function.
[1889] Task organization and reminder function
[1890] Step 1: Enter task information
[1891] User: Enter task information into Microsoft To-Do or Google Calendar.
[1892] Input: Task information entered by the user.
[1893] Output: Task information entered into the terminal.
[1894] Step 2: Submit task data
[1895] Terminal: Sends entered task information to the server in real time. Uses the HTTPS protocol.
[1896] Input: Task information.
[1897] Output: Task information sent to the server.
[1898] Step 3: Task Analysis
[1899] Server: Analyzes task data and extracts tasks with approaching deadlines.
[1900] Input: Submitted task information.
[1901] Output: Extracted tasks with approaching deadlines.
[1902] Step 4: Create a reminder notification
[1903] Server: Creates reminder notifications for tasks with approaching deadlines.
[1904] Input: Extracted tasks with approaching deadlines.
[1905] Output: Reminder notification.
[1906] Step 5: Displaying reminder notifications
[1907] Terminal: Displays the created reminder notification to the user at the appropriate time.
[1908] Input: Reminder notification.
[1909] Output: Reminder notification displayed to the user.
[1910] Step 6: Task Management
[1911] User: Check reminder notifications and manage tasks.
[1912] Input: The displayed reminder notification.
[1913] Output: Managed tasks.
[1914] The above steps establish the processing flow of this system. This allows users to efficiently manage and perform various tasks and daily routines.
[1915] (Application Example 1)
[1916] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1917] In modern security operations, there is a demand for efficient information collection and management, as well as rapid response. However, manual reporting during patrols is time-consuming and laborious, and prone to information omissions and communication errors. Furthermore, daily tasks such as task reminders and initial email replies can add to the workload, reducing overall efficiency. In this context, there is a growing need for systems that convert speech to text in real time, automatically generate responses, and provide task reminders.
[1918] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1919] In this invention, the server includes means for collecting audio data through the microphone of a terminal used by the user, means for transmitting the collected audio data to the server, means for the server to analyze the audio data and convert it into text, means for generating an appropriate response sentence using a generation AI model based on the converted text, means for displaying the generated response sentence on the terminal in real time, means for transmitting task data registered by the user to the server, means for the server to analyze the task data and create a reminder notification for tasks with approaching deadlines, and means for displaying the reminder notification on the terminal. This enables real-time information gathering and response in security operations, thereby improving operational efficiency.
[1920] "User" refers to an individual or group that uses the system.
[1921] "Terminal" refers to a digital device used by a user (e.g., a smartphone or PC).
[1922] A "microphone" refers to a device that converts sound into electrical signals.
[1923] "Audio data" refers to the digital signal of sound collected through a microphone.
[1924] A "server" refers to a computer system that receives audio data, analyzes it, and converts it into text.
[1925] A "generative AI model" refers to artificial intelligence that generates appropriate responses or suggestions based on input text data.
[1926] "Text" refers to written information converted from audio data.
[1927] "Response sentence" refers to a dialogue-style reply sentence generated by a generative AI model.
[1928] "Task data" refers to information about tasks and schedules registered by the user.
[1929] A "reminder notification" refers to a notification that informs the user of tasks that are nearing their deadline.
[1930] This invention is a system for streamlining information gathering and task management in security operations. The system includes means for collecting voice data through the microphone of a terminal used by the user and transmitting the collected voice data to a server. The server analyzes the voice data, converts it to text, generates appropriate response sentences using a generative AI model, and displays them on the terminal in real time. Furthermore, the system transmits task data registered by the user to the server, which analyzes the task data, creates reminder notifications for tasks with approaching deadlines, and displays these on the terminal.
[1931] The server uses the Google Cloud Speech-to-Text API to convert speech data into text and generates response sentences based on the text data using a generative AI model (e.g., GPT-4). The terminal displays the transcribed meeting content and generated response sentences to the user in real time. This allows users to quickly and accurately obtain information and manage tasks during security work.
[1932] For example, if a night shift security staff member uses a mobile device to make a voice report during patrol, the voice is automatically converted to text and sent to the administrator in real time. Furthermore, reminder notifications are displayed based on a security checklist, helping to prevent missed tasks. This is expected to improve work efficiency and reduce errors.
[1933] The system of the present invention is specifically implemented through the following program processing. The program uses the Google Cloud Speech-to-Text API to convert speech data into text and generates response sentences using a generative AI model. Furthermore, a task reminder function is implemented using a scheduling library.
[1934] For example, the following prompt statements are possible:
[1935] "You are the designer of an application that helps streamline security services. It will transcribe voice data into text in real time and use AI to provide users with appropriate replies and task notifications. Please design an application with the following features:
[1936] Audio data collection and real-time text conversion
[1937] Generating appropriate automated replies from text data
[1938] "Recurring task reminder notifications"
[1939] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1940] Step 1:
[1941] The device collects audio data. The device collects the voice spoken by the user through the microphone in real time and saves it as data. In this step, the input is the user's voice, and the output is digitized audio data.
[1942] Step 2:
[1943] The terminal sends the collected audio data to the server. The collected audio data is uploaded to the server in real time. The input for this step is digitized audio data, and the output is audio data stored on the server.
[1944] Step 3:
[1945] The server analyzes the audio data and converts it to text. The server uses the Google Cloud Speech-to-Text API to convert the audio data to text data. The input for this step is audio data, and the output is text data. As a concrete example of how the audio analysis process works, an audio file is input to the API, and the result is returned as text.
[1946] Step 4:
[1947] The server generates an appropriate response sentence using a generative AI model based on the converted text. The generative AI model (e.g., GPT-4) is input with text data to generate a response sentence. The input for this step is text data, and the output is the generated response sentence. The generative AI model analyzes the text data to produce the most appropriate response or suggestion.
[1948] Step 5:
[1949] The server sends the generated response message to the terminal in real time. The generated response message is sent to the terminal and displayed to the user immediately. The input for this step is the response message, and the output is the response message displayed on the terminal.
[1950] Step 6:
[1951] The user registers task data on their device, and the device sends that data to the server. When a user registers a task using a task management app, that data is sent to the server in real time. In this step, the input is the task data entered by the user, and the output is the task data stored on the server.
[1952] Step 7:
[1953] The server analyzes task data and creates reminder notifications for tasks with approaching deadlines. It analyzes task data, extracts tasks with imminent deadlines, and generates reminder notifications. The input for this step is task data, and the output is the generated reminder notifications.
[1954] Step 8:
[1955] The server sends a reminder notification to the device, and the device displays it to the user. The server sends a reminder notification to the device and notifies the user. The input for this step is the reminder notification, and the output is the reminder notification displayed on the device.
[1956] The above processing steps enable real-time text conversion of audio data, generation of response statements, task management, and reminder notifications, thereby streamlining security operations.
[1957] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1958] This invention relates to a resident generative AI system for streamlining information retrieval, text conversion, task management, email correspondence, function input, and user emotion recognition and response in users' work and daily lives. This invention comprises the following components and their operation.
[1959] System Configuration
[1960] Real-time transcription of meeting content
[1961] Collection and transmission of audio data
[1962] Terminal: Collects audio during meetings in real time via the microphone.
[1963] Terminal: Digitizes the collected audio data and sends it to the server.
[1964] Audio data analysis and text conversion
[1965] Server: Inputs received audio data into a speech recognition model and converts it into text.
[1966] Server: Stores and filters the text-based data.
[1967] Server: Sends filtered text to the terminal.
[1968] Display text
[1969] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[1970] User: Review the displayed text and edit or save it as needed.
[1971] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[1972] Notification of the meaning of unfamiliar words
[1973] Text input and analysis
[1974] User: Enter text into Excel sheets or documents.
[1975] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[1976] Terminal: Sends detected words and phrases to the server.
[1977] Word meaning search
[1978] Server: Searches for the meaning of words and phrases, as well as related information.
[1979] Server: Create search results as a pop-up notification.
[1980] Meaning notification
[1981] Device: Displays a pop-up notification to the user.
[1982] User: Review the displayed explanation and access any further information you need.
[1983] Specific example: If a user doesn't understand the meaning of a particular function in Excel, selecting that function will display its meaning in a pop-up on the device. This allows the user to understand the meaning without interrupting their work.
[1984] Suggestion for an automated reply
[1985] Receiving and analyzing emails
[1986] Terminal: Notifies the user that a new email has been received.
[1987] Terminal: Sends the content of received emails to the server.
[1988] Reply generation and display
[1989] Server: Analyzes the content of the email and generates an initial reply.
[1990] Server: Sends the generated initial reply message to the terminal.
[1991] Terminal: Displays the initial reply as a draft.
[1992] User: Review the draft, make any necessary corrections, and then send a reply.
[1993] Specific example: When a user receives an email, the server analyzes its contents and generates an initial reply such as, "Could you please schedule a meeting on the following dates?" The user can then review and send the email, allowing for a quick response.
[1994] Function suggestions for Excel and spreadsheets
[1995] Text input and analysis
[1996] User: Enter data into Excel or a spreadsheet.
[1997] Terminal: Analyzes the input data and requests function suggestions from the server.
[1998] Function proposal generation and input assistance
[1999] Server: Selects the most appropriate function or script based on the received data.
[2000] Server: Generates the selected functions and scripts as suggestions and sends them to the terminal.
[2001] Terminal: Automatically inserts the suggested function into the cell.
[2002] User: Review the proposed function and set the required data range and conditions.
[2003] Specific example: When a user aggregates sales data, the "SUM function" is suggested, and the terminal automatically inserts the function. The user can perform the aggregation simply by specifying the data range.
[2004] Task organization and reminder function
[2005] Task data collection and transmission
[2006] User: Enter tasks into a task management app or scheduler.
[2007] Terminal: Sends entered task data to the server in real time.
[2008] Task analysis and reminders
[2009] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[2010] Server: Generates reminder notifications for identified tasks.
[2011] Server: Sends a reminder notification to the device.
[2012] Display reminder notifications
[2013] Device: Displays reminder notifications to the user at the appropriate time.
[2014] User: Check the reminder notification and take action on the task.
[2015] Specific example: When a user uses a task management app and the deadline approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user can then see the reminder and efficiently work on the task.
[2016] Additional configuration for the emotion engine
[2017] Recognition and response to emotions
[2018] Collection and analysis of voice and input data
[2019] Device: Collects voice and text input in real time.
[2020] Terminal: Sends collected data to the server.
[2021] Recognition of emotions
[2022] Server: Uses an emotion engine to analyze the user's emotional state from collected audio and text data.
[2023] Server: Classifies the user's emotional state based on the analysis results and determines appropriate feedback and responses.
[2024] Feedback and implementation of countermeasures
[2025] Server: Generates feedback and responses based on the user's emotional state and sends them to the terminal.
[2026] Terminal: Displays generated feedback and responses to the user.
[2027] Example 1: When a user receives an emotionally charged email, the emotion engine detects the user's stress level and suggests softening the tone of the reply. The device displays the suggestion, and the user can send it as is.
[2028] Example 2: If a user is working with a spreadsheet and the emotion engine detects the user's confusion or stress, the device will display appropriate guidance and support to help the user.
[2029] Specific example 3: In user task management, if the emotion engine detects user fatigue or anxiety, the server adjusts task priorities and sends a reminder notification to the device. The user can then check the notification and efficiently proceed with their tasks.
[2030] By combining the system of the present invention with an emotion engine, it becomes possible to recognize the user's emotional state and provide appropriate feedback and support. This allows users to work more comfortably and efficiently in their professional and daily lives.
[2031] The following describes the processing flow.
[2032] Additional configuration for the emotion engine
[2033] Recognition and response to emotions
[2034] Collection and analysis of voice and input data
[2035] Step 1:
[2036] Terminal: Collects audio from the user's voice through the microphone.
[2037] Step 2:
[2038] Terminal: Collects text data entered by the user.
[2039] Step 3:
[2040] Terminal: Sends collected audio and text data to the server.
[2041] Recognition of emotions
[2042] Step 1:
[2043] Server: Analyzes the collected audio data and determines the user's emotional state from the audio data.
[2044] Step 2:
[2045] Server: Analyzes collected text data and uses that data to analyze the user's emotional state.
[2046] Step 3:
[2047] Server: Integrates the analysis results of voice and text data to classify the user's overall emotional state.
[2048] Feedback and implementation of countermeasures
[2049] Step 1:
[2050] Server: Determines appropriate feedback and support based on emotional state.
[2051] Step 2:
[2052] Server: Generates the determined feedback and support details and sends them to the terminal.
[2053] Step 3:
[2054] Terminal: Displays generated feedback and support details to the user.
[2055] Specific example
[2056] Example 1: Emotional email responses
[2057] Step 1:
[2058] Terminal: Collects the content of emails received by the user and sends it to the server.
[2059] Step 2:
[2060] Server: Analyzes email content and uses an emotion engine to analyze the user's emotional state.
[2061] Step 3:
[2062] Server: Based on the analysis results, it generates a reply message to alleviate user stress.
[2063] Step 4:
[2064] Terminal: Displays the generated reply as a draft to the user.
[2065] Step 5:
[2066] User: Review the draft, make any necessary corrections, and then send a reply.
[2067] Example 2: Spreadsheet operation support
[2068] Step 1:
[2069] Terminal: Collects data entered by the user in a spreadsheet and sends it to the server.
[2070] Step 2:
[2071] Server: Uses an emotion engine to analyze the confusion and stress the user is experiencing.
[2072] Step 3:
[2073] Server: Generates appropriate guides and support content based on the analysis results.
[2074] Step 4:
[2075] Terminal: Displays generated guides and support information to the user.
[2076] Specific example 3: Adjusting task management and reminders
[2077] Step 1:
[2078] Terminal: Collects task data entered by the user into the task management app and sends it to the server.
[2079] Step 2:
[2080] Server: Uses an emotion engine to analyze user fatigue and anxiety.
[2081] Step 3:
[2082] Server: Based on the analysis results, adjust task priorities and reminder timings.
[2083] Step 4:
[2084] Server: Generates a coordinated reminder notification and sends it to the device.
[2085] Step 5:
[2086] Terminal: Displays the generated reminder notification to the user at the appropriate time.
[2087] Step 6:
[2088] User: Check the displayed reminder notification and take action on the task.
[2089] Thus, by combining an emotion engine, the system of the present invention can recognize the user's emotional state in real time and provide appropriate feedback and support, thereby improving efficiency in work and daily life.
[2090] (Example 2)
[2091] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[2092] Modern information processing devices and various user systems require the efficient, real-time processing of a wide variety of tasks. However, centralized systems for quickly and accurately performing tasks such as converting audio data to text, searching for the meaning of unknown words and phrases, and automatically replying to emails are still not adequately developed. The lack of such systems can lead to decreased work efficiency and increased stress for users.
[2093] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing audio data and converting it into text using a speech recognition model, means for using the internet to search for the meaning of words and phrases and related information, and means for analyzing the content of emails and generating a primary reply using a generation AI model. This enables users to quickly record meeting content, instantly understand the meaning of unfamiliar words, and efficiently reply to emails.
[2094] "Audio data" refers to information that represents audio from meetings, conversations, etc., in a digital format.
[2095] "Digitalization" is the process of converting analog signals into digital signals.
[2096] A "server" is a computer that provides services and data to multiple clients over a network.
[2097] A "speech recognition model" is a machine learning model used to convert speech data into text.
[2098] "Text conversion" is the process of converting data such as audio or handwritten notes into text data.
[2099] "Filtering" is the process of applying specific conditions to data to extract only the necessary information.
[2100] An "information processing device" is an electronic computer that has the functions of inputting, processing, and outputting data.
[2101] An "unknown word or phrase" is a word or part of a sentence that the user cannot understand or find unclear.
[2102] A "pop-up notification" is a notification window that temporarily appears on the user's screen.
[2103] "Email" refers to messages that are sent and received electronically via the internet.
[2104] A "generative AI model" is a machine learning model used to generate text and data using artificial intelligence.
[2105] A "primary reply" is the initial response suggested in an email reply.
[2106] A "draft" is a preliminary version of a document or email before any revisions are made.
[2107] This invention relates to a resident generative AI system for streamlining information retrieval, text conversion, task management, email correspondence, function input, and user emotion recognition and response in users' work and daily lives. This invention comprises the following components and their operation.
[2108] Real-time transcription of meeting content
[2109] This system collects audio data in real time during meetings through the microphone of the user's information processing device, digitizes it, and sends it to a server. The server converts the received audio data into text using a speech recognition model (e.g., Google Cloud Speech-to-Text API). The transcribed data is stored in a database, where important keywords are filtered. The filtered text is sent to the information processing device in real time and displayed in a dedicated window.
[2110] Specific example:
[2111] When a user is conducting an online meeting, the information processing device collects audio data through its microphone and sends the digitized data to a server. The server uses a speech recognition model to transcribe the data into text, filters it, and then sends it back to the information processing device, allowing the meeting content to be displayed as text in real time.
[2112] Example of a prompt:
[2113] "Please explain, step by step, the process of the system that transcribes meeting content into text in real time."
[2114] Notification of the meaning of unfamiliar words
[2115] The system analyzes the text entered by the user into the information processing device in real time to detect unknown words and phrases. These detected words and phrases are sent to a server. The server searches for the meaning and related information of these words and phrases using the internet (e.g., Wikipedia API, Oxford Dictionaries API) and displays it on the information processing device as a pop-up notification.
[2116] Specific example:
[2117] When a user enters a specific function into an Excel sheet, it is detected and sent to the server. The server uses the corresponding API to look up the meaning of the function and displays it on the information processing device as a pop-up notification.
[2118] Example of a prompt:
[2119] "Please explain, step by step, the process of a system that automatically notifies you of the meaning of unknown words in Excel."
[2120] Suggestion for an automated reply
[2121] The server analyzes the content of the received email and generates a preliminary reply using a generative AI model (e.g., OpenAI GPT-3). The generated preliminary reply is displayed as a draft on the information processing device, and the user reviews it, makes corrections, and then sends it.
[2122] Specific example:
[2123] When a user receives an email, the server analyzes its contents and generates a preliminary reply such as, "Would you be able to schedule a meeting on the following dates?" The user can then review, revise, and send the reply, which is displayed as a draft.
[2124] Example of a prompt:
[2125] "Please explain the process of the system that generates automated reply messages, step by step."
[2126] Task organization and reminder function
[2127] When a user enters a task into a task management app or scheduler, that task data is sent to the server in real time. The server analyzes the task data, identifies unattended tasks with deadlines, and generates reminder notifications. These reminder notifications are sent to an information processing device and displayed to the user at the appropriate time.
[2128] Specific example:
[2129] When a user uses a task management app and the deadline approaches, the information processing device notifies them with a message such as, "Please complete this task by 3 PM tomorrow." The user can then use the reminder to efficiently complete the task.
[2130] Example of a prompt:
[2131] "Please describe the process of the system that sends task reminder notifications, step by step."
[2132] Recognition and response to emotions
[2133] This system incorporates an emotion engine that analyzes the user's emotional state using voice and text data collected from the user. This allows the system to classify the user's emotional state and determine appropriate feedback and responses. The generated feedback and responses are displayed on the information processing device.
[2134] Specific example:
[2135] When a user receives an emotionally charged email, the emotion engine detects the user's stress and suggests softening the tone of the reply. If the emotion engine detects confusion or stress while a user is working with a spreadsheet, the device displays appropriate guidance and support. In task management, if the emotion engine detects user fatigue or anxiety, the server adjusts task priorities and sends reminder notifications to the information processing device.
[2136] Example of a prompt:
[2137] "Please describe, step by step, how the system recognizes and appropriately responds to user emotions."
[2138] The flow of the specific processing in Example 2 will be explained using Figure 13.
[2139] Real-time transcription of meeting content
[2140] Step 1:
[2141] Voice collection
[2142] Device: Collects audio during meetings in real time using a microphone.
[2143] Input: Analog audio from a meeting.
[2144] Output: Analog audio data.
[2145] Specific operation: By pressing the "Start Recording" button, the device activates the microphone and begins collecting audio data.
[2146] Step 2:
[2147] Audio digitization
[2148] Terminal: Digitizes the collected analog audio data.
[2149] Input: Analog audio data.
[2150] Output: Digital audio data.
[2151] Specific operation: Digitize the audio at a sampling rate of 44.1kHz and convert it to PCM format.
[2152] Step 3:
[2153] Sending audio data
[2154] Terminal: Sends digitized audio data to the server.
[2155] Input: Digital audio data.
[2156] Output: Transmission of digital audio data to the server.
[2157] Specific operation: Digitized audio data is sent to the server using a real-time streaming protocol (e.g., WebSocket).
[2158] Step 4:
[2159] Receiving audio data
[2160] Server: The server receives audio data in real time.
[2161] Input: Digital audio data transmitted from the device.
[2162] Output: Audio data stored in the buffer.
[2163] Specific operation: The server continuously receives frames of audio data and stores them in a fixed buffer.
[2164] Step 5:
[2165] Speech recognition and text conversion
[2166] Server: Converts received audio data into text using a speech recognition model.
[2167] Input: Audio data stored in the buffer.
[2168] Output: Text data.
[2169] Specific operation: When the audio data buffer reaches a certain amount, the data is sent to the speech recognition API as a request, and the text result is received.
[2170] Step 6:
[2171] Text filtering and saving
[2172] Server: Stores the digitized data in a database and filters for important keywords.
[2173] Input: Text data.
[2174] Output: Filtered text data.
[2175] Specific operation: When saving text to the database, an SQL query is used to set the "importance" field, and if a specific keyword is included, it is set to "high," etc.
[2176] Step 7:
[2177] Sending filtered text
[2178] Server: Transmits text data to the information processing device in real time.
[2179] Input: Filtered text data.
[2180] Output: Sending text data to an information processing device.
[2181] Specific operation: The server sends filtered text data to the terminal via WebSocket.
[2182] Step 8:
[2183] Display text
[2184] Terminal: Displays text data in real time in a dedicated window.
[2185] Input: Filtered text data.
[2186] Output: The text displayed to the user.
[2187] Specific operation: Uses Javascript and other front-end technologies to manipulate the DOM within the window and render text.
[2188] Step 9:
[2189] Edit and save text
[2190] User: Review the displayed text, edit it as needed, and save it.
[2191] Input: The displayed text.
[2192] Output: Saved text data.
[2193] Specific operation: A text editor is provided on the editing screen, and the data is saved to local storage or cloud storage by pressing the "Save" button.
[2194] Notification of the meaning of unfamiliar words
[2195] Step 1:
[2196] Text input
[2197] User: Enter text into Excel or other documents.
[2198] Input: Text entered by the user.
[2199] Output: Input text data.
[2200] Specific action: The user enters formulas and explanations into an Excel sheet.
[2201] Step 2:
[2202] Real-time analysis
[2203] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[2204] Input: The entered text data.
[2205] Output: Unknown words or phrases detected.
[2206] Specific operation: Every time text is entered, it is analyzed using NLP (e.g., spaCy).
[2207] Step 3:
[2208] Sending to the server
[2209] Terminal: Sends detected words and phrases to the server.
[2210] Input: Unknown words or phrases detected.
[2211] Output: Data sent to the server.
[2212] Specific operation: Encode the detected data into JSON format and send it to the server via an HTTP request.
[2213] Step 4:
[2214] Search for meaning
[2215] Server: The server searches for the meaning of words and phrases using the internet.
[2216] Input: Unknown words or phrases detected.
[2217] Output: Search results data.
[2218] Specific operation: Send requests to the Wikipedia API and other dictionary APIs to retrieve meanings and related information.
[2219] Step 5:
[2220] Pop-up notification generation
[2221] Server: Generates a popup notification based on the search results.
[2222] Input: Search results data.
[2223] Output: Popup notification data.
[2224] Specific operation: Format the retrieved search results and generate a popup notification using HTML and CSS.
[2225] Step 6:
[2226] Send to device
[2227] Server: Sends a pop-up notification to the device.
[2228] Input: Pop-up notification data.
[2229] Output: Data sent to the terminal.
[2230] Specific action: The created popup notification is sent to the device as an HTTP response in JSON format.
[2231] Step 7:
[2232] Pop-up display
[2233] Device: Displays a pop-up notification on the user's screen.
[2234] Input: Pop-up notification data from the server.
[2235] Output: A pop-up notification displayed on the screen.
[2236] Specific operation: Use JavaScript to display a popup at the appropriate location and timing.
[2237] Suggestion for an automated reply
[2238] Step 1:
[2239] Email notification
[2240] Terminal: Notifies the user that a new email has been received.
[2241] Input: Received email.
[2242] Output: Notification of incoming email.
[2243] Specific action: Notify the user via a pop-up notification that a new email has been received.
[2244] Step 2:
[2245] Sending the email content
[2246] Terminal: Sends the content of received emails to the server.
[2247] Input: Received email data.
[2248] Output: Data sent to the server.
[2249] Specific operation: Convert the email content to JSON format and send it to the server via an HTTP request.
[2250] Step 3:
[2251] Generating a reply
[2252] Server: The server analyzes the email content and generates an initial reply using a generation AI model.
[2253] Input: Email content data.
[2254] Output: Initial reply.
[2255] Specific operation: Analyze the content of the email and generate a reply using a generative AI model such as OpenAI GPT-3.
[2256] Step 4:
[2257] Sending a reply
[2258] Server: Sends the generated initial reply to the terminal as a draft.
[2259] Input: Initial reply.
[2260] Output: Data sent to the terminal.
[2261] Specific operation: Convert the reply text to JSON format and send it to the terminal as an HTTP response.
[2262] Step 5:
[2263] Displaying the reply
[2264] Terminal: Displays the initial reply as a draft.
[2265] Input: Initial response data from the server.
[2266] Output: Draft below the user.
[2267] Specific operation: Use JavaScript to display a draft of the reply message in the browser.
[2268] Step 6:
[2269] User verification and correction
[2270] User: Review the draft and make any necessary corrections.
[2271] Input: Draft of the initial reply.
[2272] Output: Revised reply.
[2273] Specific operation: The user makes modifications through a text editor in the browser.
[2274] Step 7:
[2275] Corrected and sent
[2276] User: Send the revised reply.
[2277] Input: Revised reply.
[2278] Output: Sent email.
[2279] Specific action: Press the "Send" button to send the revised reply to the recipient.
[2280] The above steps demonstrate how each function is specifically implemented. Clearly defining the specific actions and data flows performed at each processing step makes it easier to understand the overall operation of the system.
[2281] (Application Example 2)
[2282] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[2283] In modern brick-and-mortar stores, employees are expected to respond to customers quickly and accurately. However, employees may be unable to provide appropriate service due to lack of experience or information. Furthermore, understanding and responding appropriately to customer emotions is also difficult. A system is needed that allows employees to receive information in real time and provide the best possible response based on customer emotions.
[2284] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing sound and text information and recognizing the emotional state of the user, means for presenting feedback and responses based on the emotional state, means for optimizing dialogue with the customer using the emotional analysis results, and means for providing support information to a visual display device for quickly and appropriately answering customer questions. This enables employees in physical stores to respond to customers quickly and appropriately.
[2285] An "information processing device" is a device used to process, store, and transmit data, and includes smartphones, tablets, and computers.
[2286] "Sound" is a wave produced by vibrations in the air, and is an audio signal collected through an input device such as a microphone.
[2287] A "network" is a communication system that allows multiple information processing devices to send and receive data from one another.
[2288] "Textual information" refers to data obtained by analyzing audio or handwritten data and converting it into text format.
[2289] "Real-time" means that data is processed and displayed almost simultaneously, with virtually no delay.
[2290] "Emotional state" refers to the emotional state of a user as analyzed using an emotion engine, and includes emotions such as joy, anger, sadness, and surprise.
[2291] "Feedback" refers to information and advice provided based on a user's behavior and circumstances.
[2292] "Response" refers to appropriate actions or measures taken in response to a specific situation or condition.
[2293] A "visual display device" is a device used to visually display data, and includes smart glasses, displays, and head-mounted displays.
[2294] A "customer" is someone who intends to purchase or use a product or service.
[2295] "Support information" refers to information and data that are provided to help users perform specific tasks.
[2296] "Emotional analysis results" refer to data obtained after evaluating and classifying emotional states.
[2297] This invention provides an information processing system for improving the customer service capabilities of employees in physical stores. The embodiments for carrying out this invention are described in detail below.
[2298] The entire system consists of a visual display device used by the user (such as smart glasses), a microphone for collecting audio, a server for analyzing the data, and a network to connect them.
[2299] 1. Speech recognition and transcription
[2300] First, the user wears smart glasses, and questions from customers in the store are collected via a microphone. The visual display device acquires the audio and transmits it to a server over the network. The server analyzes the audio data and converts it into text using a speech recognition model.
[2301] 2. Recognition of emotional states
[2302] The server analyzes customer facial expression data acquired through visual display devices and cameras, and uses an emotion engine to recognize the customer's emotional state. Emotional states include feelings such as joy, anger, sadness, and surprise. This allows the user to understand the customer's feelings.
[2303] 3. Feedback and proposed response
[2304] The server provides appropriate feedback and responses based on the recognized emotional state of the customer. For example, if the customer is angry, it suggests a response in a calm tone. If the customer is confused, it provides detailed explanations and guidance.
[2305] 4. Provision of support information
[2306] In response to customer inquiries, the server generates appropriate answers, which are displayed in real time on the smart glasses. This allows users to instantly answer customer questions. Other relevant information and product details are also presented as needed.
[2307] These features enable employees to respond to customers quickly and accurately, leading to improved customer satisfaction.
[2308] Examples
[2309] One day, if a customer in a physical store asks, "How much does this item cost?", the microphone in the smart glasses captures the voice and transmits it to a server via the network. The server converts the voice into text and also recognizes the customer's emotions from their facial expressions. Based on the emotional state, the server provides a calm response, such as, "This item costs 299 yen." This information is displayed in real time on the smart glasses' screen, allowing employees to respond to customers quickly.
[2310] Example of a prompt
[2311] "A customer has asked a question about a product. Based on that question and the subsequent conversation, please have the AI generate an appropriate answer. Also, please suggest responses that take into account the customer's emotional state."
[2312] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[2313] Step 1:
[2314] The user wears smart glasses and receives customer questions via voice. A microphone built into the smart glasses collects the voice data. This voice data is input into the user's device (smart glasses).
[2315] Step 2:
[2316] The terminal digitizes the collected audio data and transmits it to a server via the network. Here, the data format is converted, and the audio data is input to the server.
[2317] Step 3:
[2318] The server inputs the received audio data into a speech recognition model (e.g., Google Speech-to-Text API) and converts it into text information. In this process, the audio data is processed into text data, and the result is output to the server.
[2319] Step 4:
[2320] The server analyzes the converted text information and extracts specific keywords (e.g., product name, price). Furthermore, it searches for supporting information (such as price information from a database) to generate the most suitable response based on that information. These search results are then output to the server.
[2321] Step 5:
[2322] The server uses the smart glasses' camera to acquire customer facial expression data and analyzes it in real time. Using an emotion recognition engine such as DeepFace, it classifies and recognizes the customer's emotional state (e.g., anger, joy). As a result, the emotional data is output to the server.
[2323] Step 6:
[2324] The server generates an action plan based on the recognized emotional state and the content of the question. For example, if the customer is angry, it will suggest a polite way to handle the situation. This information is then sent from the server to the terminal as feedback.
[2325] Step 7:
[2326] The device (smart glasses) displays feedback information (responses and suggested actions) sent from the server in real time. The user then uses this display to respond to or suggest to the customer.
[2327] Step 8:
[2328] When a user provides an appropriate response or suggestion to a customer, the terminal additionally records that information and sends it to the server as needed to improve the relevant data. This allows the system to learn and further improve its response in the future.
[2329] This processing flow is expected to significantly improve the customer service capabilities of employees in physical stores, leading to increased customer satisfaction.
[2330] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[2331] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2332] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[2333] [Fourth Embodiment]
[2334] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[2335] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[2336] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[2337] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[2338] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[2339] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[2340] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[2341] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[2342] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[2343] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[2344] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[2345] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[2346] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2347] This invention relates to a resident generation AI system for streamlining information retrieval, text conversion, task management, email correspondence, function input, and other tasks in users' work and daily lives. This invention comprises the following components and their operation.
[2348] System Configuration
[2349] Real-time transcription of meeting content
[2350] Collection and transmission of audio data
[2351] Terminal: Collects audio in real time via microphone during meetings and conversations.
[2352] Terminal: Digitizes the collected audio data and sends it to the server.
[2353] Audio data analysis and text conversion
[2354] Server: Inputs the received audio data into the speech recognition model and converts it into text.
[2355] Server: Temporarily stores the text-based data and provides the user with appropriately filtered content.
[2356] Display text
[2357] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[2358] User: Review the displayed text and edit or save it as needed.
[2359] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[2360] Notification of the meaning of unfamiliar words
[2361] Text input and analysis
[2362] User: Enter text into Excel sheets or documents.
[2363] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[2364] Word meaning search
[2365] Server: Searches for the meaning of a word or phrase selected by the user from related databases.
[2366] Server: Analyzes relevant information and extracts appropriate content.
[2367] Meaning notification
[2368] Device: Display search results as a pop-up notification.
[2369] User: Check the displayed explanation and quickly obtain the necessary information.
[2370] Specific example: If a user doesn't understand the meaning of a particular function in Excel, selecting that function will display its meaning in a pop-up on the device. This allows the user to understand the meaning without interrupting their work.
[2371] Suggestion for an automated reply
[2372] Receiving and analyzing emails
[2373] Terminal: Notifies the user that a new email has been received and displays its contents.
[2374] Server: Analyzes the content of the email and generates an appropriate initial reply.
[2375] Reply generation and display
[2376] Server: Generates an initial reply and applies it to the template.
[2377] Terminal: Displays the generated initial reply as a draft.
[2378] User: Review the draft, make any necessary corrections, and then send a reply.
[2379] Specific example: When a user receives an email, the server analyzes its contents and generates an initial reply such as, "Could you please schedule a meeting on the following dates?" The user can then review and send the email, allowing for a quick response.
[2380] Function suggestions for Excel and spreadsheets
[2381] Text input and analysis
[2382] User: Enter data into Excel or a spreadsheet.
[2383] Terminal: Analyzes the input data and suggests appropriate functions or scripts.
[2384] Function proposal generation and input assistance
[2385] Server: Based on the input, it selects the most appropriate function or script and generates it as a suggestion.
[2386] Terminal: Automatically inserts the suggested function into the cell.
[2387] User: Review the proposed function and set the required data range and conditions.
[2388] Specific example: When a user aggregates sales data, the "SUM function" is suggested, and the terminal automatically inserts the function. The user can perform the aggregation simply by specifying the data range.
[2389] Task organization and reminder function
[2390] Task data collection and transmission
[2391] User: Enter tasks into a task management app or scheduler.
[2392] Terminal: Sends entered task data to the server in real time.
[2393] Task analysis and reminders
[2394] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[2395] Server: Creates reminder notifications for tasks with approaching deadlines.
[2396] Display reminder notifications
[2397] Device: Displays reminder notifications to the user at the appropriate time.
[2398] User: Check the reminder notification and take action on the task.
[2399] Specific example: When a user uses a task management app and the deadline approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user can then see the reminder and efficiently work on the task.
[2400] The system of the present invention allows users to efficiently carry out various tasks in their work and daily lives, significantly reducing the time and effort required for information retrieval, data entry, task management, and other related activities.
[2401] The following describes the processing flow.
[2402] Real-time transcription of meeting content
[2403] Step 1:
[2404] Terminal: Collects audio during meetings in real time via the microphone.
[2405] Step 2:
[2406] Terminal: Digitizes the collected audio data and sends it to the server.
[2407] Step 3:
[2408] Server: Inputs received audio data into a speech recognition model and converts it into text.
[2409] Step 4:
[2410] Server: Stores and filters the text-based data.
[2411] Step 5:
[2412] Server: Sends filtered text to the terminal.
[2413] Step 6:
[2414] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[2415] Step 7:
[2416] User: Review the displayed text and edit or save it as needed.
[2417] Notification of the meaning of unfamiliar words
[2418] Step 1:
[2419] User: Enter text into Excel sheets or documents.
[2420] Step 2:
[2421] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[2422] Step 3:
[2423] Terminal: Sends detected words and phrases to the server.
[2424] Step 4:
[2425] Server: Searches for the meaning of words and phrases, as well as related information.
[2426] Step 5:
[2427] Server: Create search results as a pop-up notification.
[2428] Step 6:
[2429] Device: Displays a pop-up notification to the user.
[2430] Step 7:
[2431] User: Review the displayed explanation and access any further information you need.
[2432] Suggestion for an automated reply
[2433] Step 1:
[2434] Terminal: Notifies the user that a new email has been received.
[2435] Step 2:
[2436] Terminal: Sends the content of received emails to the server.
[2437] Step 3:
[2438] Server: Analyzes the content of the email and generates an initial reply.
[2439] Step 4:
[2440] Server: Sends the generated initial reply message to the terminal.
[2441] Step 5:
[2442] Terminal: Displays the initial reply as a draft.
[2443] Step 6:
[2444] User: Review the draft, make any necessary corrections, and then submit.
[2445] Function suggestions for Excel and spreadsheets
[2446] Step 1:
[2447] User: Enter data into Excel or a spreadsheet.
[2448] Step 2:
[2449] Terminal: Analyzes the input data and requests function suggestions from the server.
[2450] Step 3:
[2451] Server: Selects the most appropriate function or script based on the received data.
[2452] Step 4:
[2453] Server: Generates the selected functions and scripts as suggestions and sends them to the terminal.
[2454] Step 5:
[2455] Terminal: Automatically inserts the suggested function into the cell.
[2456] Step 6:
[2457] User: Review the proposed function and set the required data range and conditions.
[2458] Task organization and reminder function
[2459] Step 1:
[2460] User: Enter tasks into a task management app or scheduler.
[2461] Step 2:
[2462] Terminal: Sends entered task data to the server in real time.
[2463] Step 3:
[2464] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[2465] Step 4:
[2466] Server: Generates reminder notifications for identified tasks.
[2467] Step 5:
[2468] Server: Sends a reminder notification to the device.
[2469] Step 6:
[2470] Device: Displays reminder notifications to the user at the appropriate time.
[2471] Step 7:
[2472] User: Check the reminder notification and take action on the task.
[2473] (Example 1)
[2474] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2475] In modern work and daily life, there is a demand for efficiently managing and executing a wide range of tasks. In particular, there is a lack of tools to streamline tasks such as recording meeting content, clarifying the meaning of unfamiliar words, responding quickly to emails, analyzing complex data and proposing functions, and task reminders. Furthermore, users face the problem of having to spend a significant amount of time and effort performing these tasks individually.
[2476] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[2477] In this invention, the server includes means for analyzing audio data and converting it into text, means for searching for the meaning and related information of detected words and phrases, means for generating a preliminary reply and displaying it as a draft on the terminal, means for analyzing data and generating appropriate functions and scripts, and means for analyzing task information and creating reminder notifications. This makes it possible to efficiently manage a wide range of tasks and significantly reduce the user's working time and effort.
[2478] "Audio data" refers to data that represents audio in a digital format.
[2479] A "server" is a computer system that processes and stores data over a network.
[2480] A "device" refers to a computer, smartphone, tablet, or other device used by a user.
[2481] A "user" is an individual who operates this system to perform work-related or daily life tasks.
[2482] A "microphone" is a device that collects sound and converts it into an electrical signal.
[2483] "Analysis" is the process of breaking down and interpreting data to reveal its contents.
[2484] "Text" is a collection of information expressed in written form.
[2485] "Real-time" refers to processing that occurs almost simultaneously without delay.
[2486] "Editing" is the process of modifying or changing text or other data.
[2487] "Saving" refers to the act of storing data in a storage device.
[2488] An "unknown word" is a word or phrase whose meaning is unknown to the user or system.
[2489] "Searching" is the process of finding specific information.
[2490] A "pop-up notification" is a notification window that temporarily appears on the screen.
[2491] A "first reply" is the initial draft of a response in communication such as email.
[2492] A "draft" is a preliminary version of a document created before it becomes a formal document.
[2493] A "function" is a mathematical or programming operation that takes a number or string as input and outputs a specific result.
[2494] A "script" is a series of commands or instructions that automatically execute a specific process.
[2495] A "task" is a specific activity or action in one's work or daily life.
[2496] A "reminder notification" is a notification that reminds you of the completion or deadline of a specific task.
[2497] A "generative AI model" is an artificial intelligence model that uses machine learning techniques to generate new data.
[2498] This invention relates to a resident generation AI system for users to efficiently manage and perform a wide range of tasks in their work and daily lives. This system provides real-time text conversion of voice data, notification of the meaning of unknown words, suggestions for automatic replies, suggestions for functions and scripts, and task reminder functions. Specific embodiments of this invention are described below.
[2499] System Configuration
[2500] Real-time transcription of meeting content
[2501] 1. Collection and transmission of audio data
[2502] Terminal: Collects audio in real time via the microphone during meetings and conversations. For example, use an audio capture library (e.g., PortAudio) on the terminal's operating system.
[2503] Terminal: Digitizes the collected audio data, performs encoding, and then sends it to the server using the HTTPS protocol.
[2504] 2. Analysis of audio data and text conversion
[2505] Server: Inputs the received audio data into a speech recognition API such as Google Cloud Speech-to-Text and converts it into text data.
[2506] 3. Temporary storage and real-time display of text data
[2507] Server: Temporarily stores the converted text data in a database (e.g., MySQL).
[2508] Terminal: Displays filtered text data in a dedicated window in real time. For example, it updates the page using JavaScript and WebSocket.
[2509] User: Review the displayed text and edit or save it as needed. The edited text will be saved to your local disk or cloud storage (e.g., Google Drive).
[2510] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server using Google Cloud Speech-to-Text. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[2511] Notification of the meaning of unfamiliar words
[2512] 1. Text input and analysis
[2513] User: Enter text into Excel or Google Docs.
[2514] Terminal: Analyzes the input text in real time and detects words that do not exist in the dictionary database (e.g., Oxford Dictionary API).
[2515] 2. Word meaning search and notification
[2516] Server: Uses the received word to call the Oxford Dictionaries API and retrieve its meaning and related information.
[2517] Device: Display search results as a pop-up notification. For example, using DOM manipulation and JavaScript.
[2518] User: Check the displayed explanation and obtain the necessary information.
[2519] Specific example: If a user tries to use the SUM function in Excel and doesn't understand its meaning, simply selecting the function name will cause the terminal to display a pop-up based on information from the Oxford Dictionaries. This allows the user to understand the meaning of the function without interrupting their work.
[2520] Suggestion for an automated reply
[2521] 1. Receiving and notifying emails
[2522] Terminal: Notifies the user that a new email has been received and displays its contents in the email client (e.g., Microsoft Outlook).
[2523] 2. Content analysis and generation of initial reply
[2524] Server: Analyzes the email content using a generation AI model (e.g., OpenAI GPT-4) and generates an initial reply.
[2525] 3. Draft display and confirmation
[2526] Terminal: Displays the generated initial reply as a draft.
[2527] User: Review the draft, make any necessary corrections, and send a reply.
[2528] Specific example: When a user receives a new email, the server analyzes its contents and generates an initial reply message such as, "Can you make any necessary adjustments?" The user can then review it, make any necessary corrections, and quickly send a reply.
[2529] Function suggestions for Excel and spreadsheets
[2530] 1. Data Input and Analysis
[2531] User: Enter data into Microsoft Excel or Google Sheets.
[2532] Terminal: Analyzes input data in real time. Using the Python pandas library is recommended.
[2533] 2. Function suggestions and input assistance
[2534] Server: Selects the most appropriate function based on the data content and generates a proposal using an AI model (e.g., OpenAI).
[2535] Terminal: Automatically inserts the suggested function into the cell.
[2536] User: Review the proposed function and set the data range and conditions.
[2537] Specific example: When a user aggregates sales data in a spreadsheet, a function like "=SUM(A2:A10)" is suggested and automatically inserted. The user can then easily check the data range and perform the calculation.
[2538] Task organization and reminder function
[2539] 1. Enter and submit task information
[2540] User: Enter task information into Microsoft To-Do or Google Calendar.
[2541] Terminal: Sends entered task information to the server in real time. Uses the HTTPS protocol.
[2542] 2. Task analysis and creation of reminder notifications
[2543] Server: Analyzes task information, extracts tasks with approaching deadlines, and creates reminder notifications.
[2544] 3. Displaying reminder notifications
[2545] Device: Displays reminder notifications to the user at the appropriate time.
[2546] User: Check reminder notifications and manage tasks.
[2547] Specific example: A user uses a task management app, and as the deadline for a task approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user receives this notification and can manage their tasks appropriately.
[2548] Examples of prompts for generative AI models
[2549] 1. "I want to collect the audio from the meeting and display it as text in real time. Please convert this audio to text."
[2550] 2. "I want to understand the meaning of functions used in Excel, so please explain the SUM function."
[2551] 3. "You have received a new email. Please generate an initial reply to this email immediately."
[2552] 4. "I'm entering sales data in Google Sheets, so please suggest an appropriate function."
[2553] 5. "I have set up reminder notifications in Google Calendar. The deadline for this task is approaching, so please create a reminder notification."
[2554] This invention's system allows users to efficiently carry out various tasks in their work and daily lives. It reduces the time and effort required for information retrieval, data entry, task management, etc., and provides a more efficient work environment.
[2555] The flow of the specific processing in Example 1 will be explained using Figure 11.
[2556] Real-time transcription of meeting content
[2557] Step 1: Collect audio data
[2558] Terminal: Collects audio data via the microphone during meetings and conversations. This audio data is an analog signal and is sent from the microphone to the terminal's audio input port.
[2559] Input: Analog audio signal collected via microphone.
[2560] Output: An analog audio signal is sent to the terminal.
[2561] Step 2: Digitize and transmit audio data
[2562] Terminal: Performs an ADC (analog-to-digital converter) to convert analog audio signals into digital data. After conversion, encodes the digital audio data (e.g., MP3 or WAV) and sends it to the server using the HTTPS protocol.
[2563] Input: Analog audio signal.
[2564] Output: Digital audio data is sent to the server.
[2565] Step 3: Analyzing the audio data
[2566] Server: Inputs the received digital audio data into a speech recognition API (e.g., Google Cloud Speech-to-Text) for analysis. The API converts the audio data into text data.
[2567] Input: Digital audio data.
[2568] Output: Convert to text data.
[2569] Step 4: Save text data
[2570] Server: Temporarily stores the converted text data in a database (e.g., MySQL).
[2571] Input: Converted text data.
[2572] Output: Temporarily saved text data.
[2573] Step 5: Filter the text data.
[2574] Server: Extracts important keywords and phrases from stored text data and performs filtering to remove unnecessary noise. For example, natural language processing (NLP) techniques are used to extract content appropriate to the meeting context.
[2575] Input: Temporarily saved text data.
[2576] Output: Filtered text data.
[2577] Step 6: Real-time display of text
[2578] Terminal: Displays filtered text data in a dedicated window in real time. For example, it updates the page using JavaScript and WebSocket.
[2579] Input: Filtered text data.
[2580] Output: Text displayed in real time.
[2581] Step 7: Edit and save the text
[2582] User: Review the displayed text and edit or save it as needed. The edited text will be saved to your local disk or cloud storage (e.g., Google Drive).
[2583] Input: Text displayed in real time.
[2584] Output: Edited and saved text data.
[2585] Notification of the meaning of unfamiliar words
[2586] Step 1: Enter text
[2587] User: Enter text into Excel or Google Docs.
[2588] Input: Text entered by the user.
[2589] Output: Text is entered into the terminal.
[2590] Step 2: Detecting unknown words
[2591] Terminal: Analyzes the input text in real time and detects words that do not exist in the dictionary database (e.g., Oxford Dictionary API).
[2592] Input: The entered text.
[2593] Output: List of unknown words.
[2594] Step 3: Sending a word
[2595] Terminal: Sends unknown words to the server. For example, using an Ajax request.
[2596] Input: List of unknown words.
[2597] Output: Unknown word sent to the server.
[2598] Step 4: Semantic retrieval and analysis
[2599] Server: Uses a dictionary API to search for the meaning of the submitted word and related information, and analyzes the related information.
[2600] Input: An unknown word sent to the server.
[2601] Output: Meaning and related information.
[2602] Step 5: Notification of Meaning
[2603] Terminal: Displays the analyzed meaning and related information as a popup window. Specifically, it uses DOM manipulation and JavaScript.
[2604] Input: Meaning and related information.
[2605] Output: Meaning displayed in the pop-up notification.
[2606] Step 6: Confirm the contents
[2607] User: Check the displayed explanation and obtain the necessary information.
[2608] Input: Meaning of the pop-up notification.
[2609] Output: Information obtained by the user.
[2610] Suggestion for an automated reply
[2611] Step 1: Receiving and receiving emails
[2612] Terminal: Notifies the user that a new email has been received and displays its contents in the email client (e.g., Microsoft Outlook).
[2613] Input: Received email.
[2614] Output: Email notification displayed in the email client.
[2615] Step 2: Analyze the content of the email
[2616] Server: Sends the body of the received email to a generating AI model (e.g., OpenAI GPT-4) for analysis.
[2617] Input: The body of the received email.
[2618] Output: Analyzed email content.
[2619] Step 3: Generating the initial reply
[2620] Server: Based on the analysis results, it generates an initial reply message and applies its content to a template.
[2621] Input: Analyzed email content.
[2622] Output: Initial reply.
[2623] Step 4: Draft Display
[2624] Terminal: Displays the generated initial reply as a draft.
[2625] Input: Initial reply message.
[2626] Output: Reply displayed as a draft.
[2627] Step 5: Review and revise the draft
[2628] User: Review the draft and make any necessary corrections.
[2629] Input: Reply displayed as a draft.
[2630] Output: Revised reply text.
[2631] Step 6: Send a reply
[2632] User: Send the revised reply.
[2633] Input: Revised reply text.
[2634] Output: Sent reply email.
[2635] Function suggestions for Excel and spreadsheets
[2636] Step 1: Enter data
[2637] User: Enter data into Microsoft Excel or Google Sheets.
[2638] Input: Data entered by the user.
[2639] Output: Data entered into the terminal.
[2640] Step 2: Analysis of the month
[2641] Terminal: Analyzes input data in real time. Uses Python's pandas library, etc.
[2642] Input: The entered data.
[2643] Output: Analysis results.
[2644] Step 3: Propose a function
[2645] Server: Based on the data analysis results, it selects the optimal function and generates proposals using a generative AI model (e.g., OpenAI).
[2646] Input: Analysis results.
[2647] Output: Proposed function.
[2648] Step 4: Function display and input assistance
[2649] Terminal: Automatically inserts the suggested function into the cell.
[2650] Input: Proposed function.
[2651] Output: The function inserted into the cell.
[2652] Step 5: Verify and configure the function
[2653] User: Review the proposed function and set the required data range and conditions.
[2654] Input: The function inserted into the cell.
[2655] Output: The configured function.
[2656] Task organization and reminder function
[2657] Step 1: Enter task information
[2658] User: Enter task information into Microsoft To-Do or Google Calendar.
[2659] Input: Task information entered by the user.
[2660] Output: Task information entered into the terminal.
[2661] Step 2: Submit task data
[2662] Terminal: Sends entered task information to the server in real time. Uses the HTTPS protocol.
[2663] Input: Task information.
[2664] Output: Task information sent to the server.
[2665] Step 3: Task Analysis
[2666] Server: Analyzes task data and extracts tasks with approaching deadlines.
[2667] Input: Submitted task information.
[2668] Output: Extracted tasks with approaching deadlines.
[2669] Step 4: Create a reminder notification
[2670] Server: Creates reminder notifications for tasks with approaching deadlines.
[2671] Input: Extracted tasks with approaching deadlines.
[2672] Output: Reminder notification.
[2673] Step 5: Displaying reminder notifications
[2674] Terminal: Displays the created reminder notification to the user at the appropriate time.
[2675] Input: Reminder notification.
[2676] Output: Reminder notification displayed to the user.
[2677] Step 6: Task Management
[2678] User: Check reminder notifications and manage tasks.
[2679] Input: The displayed reminder notification.
[2680] Output: Managed tasks.
[2681] The above steps establish the processing flow of this system. This allows users to efficiently manage and perform various tasks and daily routines.
[2682] (Application Example 1)
[2683] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2684] In modern security operations, there is a demand for efficient information collection and management, as well as rapid response. However, manual reporting during patrols is time-consuming and laborious, and prone to information omissions and communication errors. Furthermore, daily tasks such as task reminders and initial email replies can add to the workload, reducing overall efficiency. In this context, there is a growing need for systems that convert speech to text in real time, automatically generate responses, and provide task reminders.
[2685] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[2686] In this invention, the server includes means for collecting audio data through the microphone of a terminal used by the user, means for transmitting the collected audio data to the server, means for the server to analyze the audio data and convert it into text, means for generating an appropriate response sentence using a generation AI model based on the converted text, means for displaying the generated response sentence on the terminal in real time, means for transmitting task data registered by the user to the server, means for the server to analyze the task data and create a reminder notification for tasks with approaching deadlines, and means for displaying the reminder notification on the terminal. This enables real-time information gathering and response in security operations, thereby improving operational efficiency.
[2687] "User" refers to an individual or group that uses the system.
[2688] "Terminal" refers to a digital device used by a user (e.g., a smartphone or PC).
[2689] A "microphone" refers to a device that converts sound into electrical signals.
[2690] "Audio data" refers to the digital signal of sound collected through a microphone.
[2691] A "server" refers to a computer system that receives audio data, analyzes it, and converts it into text.
[2692] A "generative AI model" refers to artificial intelligence that generates appropriate responses or suggestions based on input text data.
[2693] "Text" refers to written information converted from audio data.
[2694] "Response sentence" refers to a dialogue-style reply sentence generated by a generative AI model.
[2695] "Task data" refers to information about tasks and schedules registered by the user.
[2696] A "reminder notification" refers to a notification that informs the user of tasks that are nearing their deadline.
[2697] This invention is a system for streamlining information gathering and task management in security operations. The system includes means for collecting voice data through the microphone of a terminal used by the user and transmitting the collected voice data to a server. The server analyzes the voice data, converts it to text, generates appropriate response sentences using a generative AI model, and displays them on the terminal in real time. Furthermore, the system transmits task data registered by the user to the server, which analyzes the task data, creates reminder notifications for tasks with approaching deadlines, and displays these on the terminal.
[2698] The server uses the Google Cloud Speech-to-Text API to convert speech data into text and generates response sentences based on the text data using a generative AI model (e.g., GPT-4). The terminal displays the transcribed meeting content and generated response sentences to the user in real time. This allows users to quickly and accurately obtain information and manage tasks during security work.
[2699] For example, if a night shift security staff member uses a mobile device to make a voice report during patrol, the voice is automatically converted to text and sent to the administrator in real time. Furthermore, reminder notifications are displayed based on a security checklist, helping to prevent missed tasks. This is expected to improve work efficiency and reduce errors.
[2700] The system of the present invention is specifically implemented through the following program processing. The program uses the Google Cloud Speech-to-Text API to convert speech data into text and generates response sentences using a generative AI model. Furthermore, a task reminder function is implemented using a scheduling library.
[2701] For example, the following prompt statements are possible:
[2702] "You are the designer of an application that helps streamline security services. It will transcribe voice data into text in real time and use AI to provide users with appropriate replies and task notifications. Please design an application with the following features:
[2703] Audio data collection and real-time text conversion
[2704] Generating appropriate automated replies from text data
[2705] "Recurring task reminder notifications"
[2706] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[2707] Step 1:
[2708] The device collects audio data. The device collects the voice spoken by the user through the microphone in real time and saves it as data. In this step, the input is the user's voice, and the output is digitized audio data.
[2709] Step 2:
[2710] The terminal sends the collected audio data to the server. The collected audio data is uploaded to the server in real time. The input for this step is digitized audio data, and the output is audio data stored on the server.
[2711] Step 3:
[2712] The server analyzes the audio data and converts it to text. The server uses the Google Cloud Speech-to-Text API to convert the audio data to text data. The input for this step is audio data, and the output is text data. As a concrete example of how the audio analysis process works, an audio file is input to the API, and the result is returned as text.
[2713] Step 4:
[2714] The server generates an appropriate response sentence using a generative AI model based on the converted text. The generative AI model (e.g., GPT-4) is input with text data to generate a response sentence. The input for this step is text data, and the output is the generated response sentence. The generative AI model analyzes the text data to produce the most appropriate response or suggestion.
[2715] Step 5:
[2716] The server sends the generated response message to the terminal in real time. The generated response message is sent to the terminal and displayed to the user immediately. The input for this step is the response message, and the output is the response message displayed on the terminal.
[2717] Step 6:
[2718] The user registers task data on their device, and the device sends that data to the server. When a user registers a task using a task management app, that data is sent to the server in real time. In this step, the input is the task data entered by the user, and the output is the task data stored on the server.
[2719] Step 7:
[2720] The server analyzes task data and creates reminder notifications for tasks with approaching deadlines. It analyzes task data, extracts tasks with imminent deadlines, and generates reminder notifications. The input for this step is task data, and the output is the generated reminder notifications.
[2721] Step 8:
[2722] The server sends a reminder notification to the device, and the device displays it to the user. The server sends a reminder notification to the device and notifies the user. The input for this step is the reminder notification, and the output is the reminder notification displayed on the device.
[2723] The above processing steps enable real-time text conversion of audio data, generation of response statements, task management, and reminder notifications, thereby streamlining security operations.
[2724] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[2725] This invention relates to a resident generative AI system for streamlining information retrieval, text conversion, task management, email correspondence, function input, and user emotion recognition and response in users' work and daily lives. This invention comprises the following components and their operation.
[2726] System Configuration
[2727] Real-time transcription of meeting content
[2728] Collection and transmission of audio data
[2729] Terminal: Collects audio during meetings in real time via the microphone.
[2730] Terminal: Digitizes the collected audio data and sends it to the server.
[2731] Audio data analysis and text conversion
[2732] Server: Inputs received audio data into a speech recognition model and converts it into text.
[2733] Server: Stores and filters the text-based data.
[2734] Server: Sends filtered text to the terminal.
[2735] Display text
[2736] Terminal: Displays the transcribed meeting content in real time in a dedicated window.
[2737] User: Review the displayed text and edit or save it as needed.
[2738] Specific example: When a user is conducting an online meeting, their device collects audio through the microphone and sends it to a server. The server converts the audio to text and displays it on the device in real time, allowing the user to quickly record the meeting content.
[2739] Notification of the meaning of unfamiliar words
[2740] Text input and analysis
[2741] User: Enter text into Excel sheets or documents.
[2742] Terminal: Analyzes the input text in real time and detects unknown words and phrases.
[2743] Terminal: Sends detected words and phrases to the server.
[2744] Word meaning search
[2745] Server: Searches for the meaning of words and phrases, as well as related information.
[2746] Server: Create search results as a pop-up notification.
[2747] Meaning notification
[2748] Device: Displays a pop-up notification to the user.
[2749] User: Review the displayed explanation and access any further information you need.
[2750] Specific example: If a user doesn't understand the meaning of a particular function in Excel, selecting that function will display its meaning in a pop-up on the device. This allows the user to understand the meaning without interrupting their work.
[2751] Suggestion for an automated reply
[2752] Receiving and analyzing emails
[2753] Terminal: Notifies the user that a new email has been received.
[2754] Terminal: Sends the content of received emails to the server.
[2755] Reply generation and display
[2756] Server: Analyzes the content of the email and generates an initial reply.
[2757] Server: Sends the generated initial reply message to the terminal.
[2758] Terminal: Displays the initial reply as a draft.
[2759] User: Review the draft, make any necessary corrections, and then send a reply.
[2760] Specific example: When a user receives an email, the server analyzes its contents and generates an initial reply such as, "Could you please schedule a meeting on the following dates?" The user can then review and send the email, allowing for a quick response.
[2761] Function suggestions for Excel and spreadsheets
[2762] Text input and analysis
[2763] User: Enter data into Excel or a spreadsheet.
[2764] Terminal: Analyzes the input data and requests function suggestions from the server.
[2765] Function proposal generation and input assistance
[2766] Server: Selects the most appropriate function or script based on the received data.
[2767] Server: Generates the selected functions and scripts as suggestions and sends them to the terminal.
[2768] Terminal: Automatically inserts the suggested function into the cell.
[2769] User: Review the proposed function and set the required data range and conditions.
[2770] Specific example: When a user aggregates sales data, the "SUM function" is suggested, and the terminal automatically inserts the function. The user can perform the aggregation simply by specifying the data range.
[2771] Task organization and reminder function
[2772] Task data collection and transmission
[2773] User: Enter tasks into a task management app or scheduler.
[2774] Terminal: Sends entered task data to the server in real time.
[2775] Task analysis and reminders
[2776] Server: Analyzes task data and identifies uncompleted tasks with deadlines.
[2777] Server: Generates reminder notifications for identified tasks.
[2778] Server: Sends a reminder notification to the device.
[2779] Display reminder notifications
[2780] Device: Displays reminder notifications to the user at the appropriate time.
[2781] User: Check the reminder notification and take action on the task.
[2782] Specific example: When a user uses a task management app and the deadline approaches, the device notifies them with a message saying, "Please complete this task by 3 PM tomorrow." The user can then see the reminder and efficiently work on the task.
[2783] Additional configuration for the emotion engine
[2784] Recognition and response to emotions
[2785] Collection and analysis of voice and input data
[2786] Device: Collects voice and text input in real time.
[2787] Terminal: Sends collected data to the server.
[2788] Recognition of emotions
[2789] Server: Uses an emotion engine to analyze the user's emotional state from collected audio and text data.
[2790] Server: Classifies the user's emotional state based on the analysis results and determines appropriate feedback and responses.
[2791] Feedback and implementation of countermeasures
[2792] Server: Generates feedback and responses based on the user's emotional state and sends them to the terminal.
[2793] Terminal: Displays generated feedback and responses to the user.
[2794] Example 1: When a user receives an emotionally charged email, the emotion engine detects the user's stress level and suggests softening the tone of the reply. The device displays the suggestion, and the user can send it as is.
[2795] Example 2: If a user is working with a spreadsheet and the emotion engine detects the user's confusion or stress, the device will display appropriate guidance and support to help the user.
[2796] Specific example 3: In user task management, if the emotion engine detects user fatigue or anxiety, the server adjusts task priorities and sends a reminder notification to the device. The user can then check the notification and efficiently proceed with their tasks.
[2797] By combining the system of the present invention with an emotion engine, it becomes possible to recognize the user's emotional state and provide appropriate feedback and support. This allows users to work more comfortably and efficiently in their professional and daily lives.
[2798] The following describes the processing flow.
[2799] Additional configuration for the emotion engine
[2800] Recognition and response to emotions
[2801] Collection and analysis of voice and input data
[2802] Step 1:
[2803] Terminal: Collects audio from the user's voice through the microphone.
[2804] Step 2:
[2805] Terminal: Collects text data entered by the user.
[2806] Step 3:
[2807] Terminal: Sends collected audio and text data to the server.
[2808] Recognition of emotions
[2809] Step 1:
[2810] Server: Analyzes the collected audio data and determines the user's emotional state from the audio data.
[2811] Step 2:
[2812] Server: Analyzes collected text data and uses that data to analyze the user's emotional state.
[2813] Step 3:
[2814] Server: Integrates the analysis results of voice and text data to classify the user's overall emotional state.
[2815] Feedback and implementation of countermeasures
[2816] Step 1:
[2817] Server: Determines appropriate feedback and support based on emotional state.
[2818] Step 2:
[2819] Server: Generates the determined feedback and support details and sends them to the terminal.
[2820] Step 3:
[2821] Terminal: Displays generated feedback and support details to the user.
[2822] Specific example
[2823] Example 1: Emotional email responses
[2824] Step 1:
[2825] Terminal: Collects the content of emails received by the user and sends it to the server.
[2826] Step 2:
[2827] Server: Analyzes email content and uses an emotion engine to analyze the user's emotional state.
[2828] Step 3:
[2829] Server: Based on the analysis results, it generates a reply message to alleviate user stress.
[2830] Step 4:
[2831] Terminal: Displays the generated reply as a draft to the user.
[2832] Step 5:
[2833] User: Review the draft, make any necessary corrections, and then send a reply.
[2834] Example 2: Spreadsheet operation support
[2835] Step 1:
[2836] Terminal: Collects data entered by the user in a spreadsheet and sends it to the server.
[2837] Step 2:
[2838] Server: Uses an emotion engine to analyze the confusion and stress the user is experiencing.
[2839] Step 3:
[2840] Server: Generates appropriate guides and support content based on the analysis results.
[2841] Step 4:
[2842] Terminal: Displays generated guides and support information to the user.
[2843] Specific example 3: Adjusting task management and reminders
[2844] Step 1:
[2845] Terminal: Collects task data entered by the user into the task management app and sends it to the server.
[2846] Step 2:
[2847] Server: Uses an emotion engine to analyze user fatigue and anxiety.
[2848] Step 3:
[2849] Server: Based on the analysis results, adjust task priorities and reminder timings.
[2850] Step 4:
[2851] Server: Generates a coordinated reminder notification and sends it to the device.
[2852] Step 5:
[2853] Terminal: Displays the generated reminder notification to the user at the appropriate time.
[2854] Step 6:
[2855] User: Check the displayed reminder notification and take action on the task.
[2856] Thus, by combining an emotion engine, the system of the present invention can recognize the user's emotional state in real time and provide appropriate feedback and support, thereby improving efficiency in work and daily life.
[2857] (Example 2)
[2858] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2859] Modern information processing devices and various user systems require the efficient, real-time processing of a wide variety of tasks. However, centralized systems for quickly and accurately performing tasks such as converting audio data to text, searching for the meaning of unknown words and phrases, and automatically replying to emails are still not adequately developed. The lack of such systems can lead to decreased work efficiency and increased stress for users.
[2860] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing audio data and converting it into text using a speech recognition model, means for using the internet to search for the meaning of words and phrases and related information, and means for analyzing the content of emails and generating a primary reply using a generation AI model. This enables users to quickly record meeting content, instantly understand the meaning of unfamiliar words, and efficiently reply to emails.
[2861] "Audio data" refers to information that represents audio from meetings, conversations, etc., in a digital format.
[2862] "Digitalization" is the process of converting analog signals into digital signals.
[2863] A "server" is a computer that provides services and data to multiple clients over a network.
[2864] A "speech recognition model" is a machine learning model used to convert speech data into text.
[2865] "Text conversion" is the process of converting data such as audio or handwritten notes into text data.
[2866] "Filtering" is the process of applying specific conditions to data to extract only the necessary information.
[2867] An "information processing device" is an electronic computer that has the functions of inputting, processing, and outputting data.
[2868] An "unknown word or phrase" is a word or part of a sentence that the user cannot understand or find unclear.
[2869] A "pop-up notification" is a notification window that temporarily appears on the user's screen.
[2870] "Email" refers to messages that are sent and received electronically via the internet.
[2871] A "generative AI model" is a machine learning model used to generate text and data using artificial intelligence.
[2872] A "primary reply" is the initial re...
Claims
1. A means of collecting voice data through the microphone of the device used by the user, A means of sending the collected audio data to a server, A means by which the server analyzes audio data and converts it into text, A means of displaying the converted text on the terminal in real time, A system that includes this.
2. A means for analyzing text entered or selected by the user on the device to detect unknown words or phrases, The server provides a means to search for the meaning and related information of detected words and phrases, A means of displaying search results as a pop-up notification on the device, The system according to claim 1, including the following:
3. A means by which the server analyzes the content of the received email and generates an initial reply, A means of displaying the generated initial reply as a draft on the terminal, A means for users to review, revise, and submit a draft, The system according to claim 1, including the following:
4. A means to analyze the content entered by the user into a spreadsheet and suggest appropriate functions or scripts, The server provides a means for generating proposed functions and scripts, A method for automatically inserting generated functions and scripts into cells, The system according to claim 1, including the following:
5. The means by which users input tasks into schedulers or task management apps, A means of sending the input task data to the server in real time, A method for the server to analyze task data, identify uncompleted tasks with deadlines, and send reminders, A means of displaying reminder notifications to users at the appropriate time, The system according to claim 1, including the following:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A