system

An AI-powered communication management system addresses the inefficiencies and security risks of conventional call handling by using speech recognition and automated responses to manage calls effectively, ensuring only necessary calls are received and recorded.

JP2026069112APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Conventional communication systems lack effective means to manage nuisance, business, and fraud calls, leading to inefficiencies and potential security risks by not knowing caller requirements in advance, which can result in wasted time and missed important calls.

Method used

An AI-based communication management system that includes speech recognition, natural language processing, and automated response mechanisms to identify caller information, classify call requirements, and decide whether to transfer calls based on user needs, enabling efficient call management and reducing risks.

Benefits of technology

The system efficiently filters unnecessary calls, improves work efficiency, enhances security, and ensures important calls are not missed by automating responses and recording calls as needed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069112000001_ABST
    Figure 2026069112000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of receiving calls from a communication terminal, A means of obtaining and identifying sender information from a database, A speech recognition method that converts speech to text, A natural language processing tool that analyzes transcribed information and classifies requirements, A means of determining whether or not to transfer a call to the user based on the caller's requirements, A system that includes means for providing an automated response to the caller.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the call management mechanism of conventional communication terminals, there are insufficient means to deal with nuisance calls, business calls, and fraud calls. For this reason, the recipient may waste time on unnecessary calls. Also, since the requirements of the caller are not known in advance, there is a possibility of missing an important call. As a result, the efficiency of operations in individuals and enterprises decreases, and security risks also arise. It is necessary to solve this problem.

Means for Solving the Problems

[0005] This invention solves these problems by providing an AI-based communication management system equipped with specific means. The system includes means for receiving calls from communication terminals, means for obtaining and identifying caller information from a database, speech recognition means for converting speech to text, natural language processing means for analyzing the textualized information and classifying requirements, and means for deciding whether or not to transfer the call to a user based on the caller's requirements. This enables efficient management of unnecessary calls. Furthermore, functions for recording calls as needed and notifying callers using automated response means further streamline call management and reduce risks.

[0006] A "communication terminal" is an electronic device that has the function of receiving and making telephone calls.

[0007] A "database" is a structured collection of information that systematically organizes data, allowing for efficient searching and retrieval of specific data.

[0008] "Speech recognition means" refers to a technology or device that converts speech into text data.

[0009] "Natural language processing methods" are technologies that analyze text data and perform classification and summarization based on its content.

[0010] An "automatic response mechanism" is a function in which the system automatically sends a predetermined message to the caller. [Brief explanation of the drawing]

[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4]This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0013] First, let's explain the terminology used in the following explanation.

[0014] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0015] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0016] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0017] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0019] [First Embodiment]

[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0032] This invention is an AI-powered call management system designed to efficiently receive only the calls that the user needs. When the server receives a call from a communication terminal, it retrieves caller information from a database and identifies whether the caller is an acquaintance or not. If the caller is not an acquaintance, the server, through AI, confirms the requirements with the caller.

[0033] The AI ​​uses speech recognition technology to convert speech information into text. The converted text is then analyzed using natural language processing and classified as a sales call, a scam call, or for other purposes. Based on this classification information, the server decides whether to transfer the call to the user.

[0034] For example, if the server determines that a call is a sales call, it will use an automated response function to send a rejection message to the caller. On the other hand, if the caller is an acquaintance and the server recognizes that the matter is important, the server will transfer the call directly to the user's terminal.

[0035] Furthermore, users can optionally enable a recording function, which notifies the caller that the call is being recorded. The recording data is stored on the server for later reference.

[0036] This system is useful not only for personal use but also as a business telephone system for companies. By automatically processing unnecessary calls, it can improve work efficiency and enhance security.

[0037] The following describes the processing flow.

[0038] Step 1:

[0039] The server receives a call from a communication terminal. It detects the caller's phone number and retrieves the caller's information by searching the database based on that number.

[0040] Step 2:

[0041] Based on the caller information obtained by the server, the system checks if the caller is someone the user knows. If they are, the call is forwarded to the user's device.

[0042] Step 3:

[0043] The server connects the call to the AI ​​and asks the caller about their purpose. The AI ​​receives the caller's voice and converts it into text using speech recognition technology.

[0044] Step 4:

[0045] The server receives the message information transcribed into text from the AI ​​and performs natural language processing to analyze the message. This analysis classifies the message as sales, fraud, or other.

[0046] Step 5:

[0047] The server determines whether to transfer the call to the user based on the classification results. If it is determined to be an important call, it is transferred to the user's terminal so that the user can receive the call.

[0048] Step 6:

[0049] If the server determines that a call may be a sales call or a scam, it will use an automated response function to reject the caller with a pre-set message. If the recording option is enabled, it will notify the caller and record and save the conversation.

[0050] (Example 1)

[0051] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0052] In modern society, unnecessary and nuisance calls are frequent occurrences, often hindering the efficient operation of businesses and individuals. This can lead to missed important communications and the waste of time and resources. Therefore, there is a need for systems that efficiently receive only necessary calls and appropriately handle unnecessary ones.

[0053] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0054] In this invention, the server includes means for receiving calls from a communication device, means for acquiring and identifying caller information from an information storage device, and means for speech recognition that converts speech into text. This makes it possible for users to efficiently receive only important calls without being bothered by unnecessary calls.

[0055] "Communication equipment" is a general term for hardware or software used to send and receive data.

[0056] "Caller information" refers to the identification information of the caller used when making a phone call, and usually includes a phone number.

[0057] An "information storage device" is a database or storage system that stores data and allows it to be freely accessed as needed.

[0058] "Speech recognition means" is a general term for technologies or systems that convert speech input into string-formatted data.

[0059] "Stringified information" refers to text data converted by speech recognition technology.

[0060] "Natural language processing means" is a general term for systems that analyze string-based information to understand or classify its content.

[0061] "Recording function" refers to the process or function for saving call content as digital data.

[0062] "Means of automated response" refers to a system or process that mechanically responds to calls using pre-set messages.

[0063] This invention is a system that uses AI technology to streamline communication management. It primarily involves the coordinated operation of server, terminal, and user elements.

[0064] server

[0065] The server receives calls from communication devices and retrieves caller information from an information storage device. The server uses a database system to identify the caller information. Relational databases such as MySQL (registered trademark) are one option. Based on the received call information, the server processes the voice data.

[0066] Speech recognition means

[0067] The server uses speech recognition to convert the incoming speech into text. This can be done using commercial speech recognition software or cloud-based APIs. For example, Google® Cloud Speech-to-Text can fulfill this role.

[0068] Natural language processing means

[0069] The converted string is analyzed by a natural language processing system. Here, a generative AI model is used to classify requirements based on the text data. By using the OpenAI® GPT model, it is possible to identify potential sales calls and scams with high accuracy.

[0070] Automatic response and recording function

[0071] Depending on the caller's requirements, the server will automatically respond. If it is determined to be a sales call, the server will automatically send a rejection message. Furthermore, users can enable the call recording function and save the recording to the server for later review.

[0072] Specific example

[0073] Suppose a user is in a coffee shop when they receive a call from an unknown number. In this case, the server quickly identifies it as a sales call and uses an automated response system to send a polite refusal message to the caller.

[0074] Example of a prompt

[0075] An example of a prompt for a generating AI model is: "When a call is received from an unknown caller, explain how to classify the call and notify the user. Also, specify how to handle sales calls."

[0076] This system offers users the advantage of receiving only necessary calls efficiently and securely, and also makes recording and managing calls easier.

[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0078] Step 1:

[0079] The server receives a call from a communication device. At this time, it receives the call metadata (caller ID, time, date, etc.) as input and prepares to process the call. As output, the call information is recorded in the internal system and made available as input data for the next step.

[0080] Step 2:

[0081] The server retrieves and identifies source information from the information storage device. It uses the source number as input and accesses the database via an SQL query. Specifically, the server searches for records matching the source number, and the output is the detailed information of the corresponding record.

[0082] Step 3:

[0083] The server identifies whether the sender is an acquaintance. Based on information retrieved from the database, it compares it with the user's contact list and uses the sender information as input. If the sender is not an acquaintance, it outputs "Unknown sender" and proceeds to the next step.

[0084] Step 4:

[0085] For callers who are not acquainted, the server uses speech recognition to confirm the requirements. In this process, an automated voice message is played to the caller, specifically the message, "Please state your request." The input is the caller's voice, which is converted into text using speech recognition technology. The output is the caller's requirements in text format.

[0086] Step 5:

[0087] The server analyzes the stringified information using natural language processing techniques and classifies the requirements. The input is text data, which is processed by a generative AI model. Specifically, it classifies the text data into sales calls, fraudulent calls, and other categories. The output, as the classification result, is then passed on to the next processing step.

[0088] Step 6:

[0089] Based on the classification results, the server decides whether to transfer the call to the user or to reject it with an automated response. The classification results are used as input; specifically, if it's determined to be a sales call, a message saying "We are currently unable to connect your call" is sent to the caller. If it's identified as an important matter, the call is transferred to the user as output.

[0090] Step 7:

[0091] If a user wishes to record a call, the server enables the recording function. When recording begins, the caller is notified that "this call is being recorded." Specifically, the recorded data is saved to the server as output and can be accessed by the user later.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] In today's communication environment, many unnecessary calls hinder the productivity of individuals and businesses. This is particularly true for customer support on e-commerce sites, where the wide range of customer inquiries makes it difficult to efficiently handle important calls. This can lead to delays in responding to critical inquiries that require a quick response.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for receiving calls from communication devices, means for obtaining and identifying caller information from a storage device, and means for speech recognition for converting speech into text. This makes it possible to automatically classify customer inquiries and quickly and efficiently transfer important inquiries to human operators.

[0097] "Communication equipment" refers to electronic devices used to send and receive voice and data over long distances.

[0098] A "call" is a request for communication made to another device or user using a communication device.

[0099] "Sender information" refers to data that includes the identifier and related information of the party that initiated the communication.

[0100] A "storage device" is a device that stores information and allows it to be reused as needed.

[0101] "Identification" is the process of recognizing a specific object and determining what it is.

[0102] "Speech recognition means" refers to technology that analyzes speech signals and converts them into text format.

[0103] A "text" is a composition in which meaningful words are arranged in a certain order.

[0104] "Natural language processing methods" are technologies that analyze text-based language data and understand its content and meaning.

[0105] "Classification" is the process of grouping information according to specific criteria.

[0106] "Purpose" refers to the goal set for an action or process.

[0107] A "user" is a person who uses or operates a system or device.

[0108] "Automated response" refers to a reaction or reply performed autonomously by a machine or software.

[0109] An "operator" refers to a person responsible for operating a system or piece of equipment.

[0110] In the system that implements this invention, the server is responsible for receiving calls from communication devices and identifying the caller by retrieving their information from storage. The hardware uses smartphones or dedicated terminals as communication devices, and the software uses the Python speech_recognition library as a speech recognition tool. The received speech is then converted into text. The converted text is analyzed using a natural language processing library such as TextBlob and classified according to its purpose.

[0111] Based on the analysis results, the server identifies important inquiries and quickly forwards them to human operators. This significantly improves the efficiency of customer support on e-commerce sites. For example, if a customer inquires about product shipping, it is deemed important and forwarded to an operator. At that time, a generative AI model is used to generate a prompt such as, "Convert the voice data received by customer support into text, classify the content of the inquiry, and take appropriate action."

[0112] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0113] Step 1:

[0114] The server receives a call from a communication device. At this point, the input is a call signal generated by the communication device, and the server takes this signal and outputs it as call information. Specifically, the server receives call information from the network using a communication protocol.

[0115] Step 2:

[0116] The server retrieves caller information from its storage device and identifies it. At this point, the input is the previously outputted call information, and a database query is generated based on the caller ID within it. This retrieves caller-related information from the database. This information is then provided as output. In its specific operation, SQL queries are used to retrieve the necessary data from the database.

[0117] Step 3:

[0118] The server uses speech recognition to convert received audio into text. The input is audio data, and by applying a speech recognition algorithm to it, the output is generated as text data. Specifically, the Python `speech_recognition` library is used for this process.

[0119] Step 4:

[0120] The server uses natural language processing to analyze the converted text and classify its purpose. The input is the text data generated in the previous step, and the output is the analyzed information classified by purpose. This process uses the TextBlob library for language analysis and classification.

[0121] Step 5:

[0122] The server determines, based on the classified information, whether to forward important inquiries to a human operator or to provide an automated response. The input is the classified information, and the output is the forwarding or response action. Specifically, important inquiries are forwarded to the operator's terminal using the API, while other inquiries generate automated response messages.

[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0124] This invention is an advanced call management system using AI that optimizes calls received by users by analyzing various information, including the caller's emotions. When the server receives a call from a communication terminal, it compares the caller's number information with a database to identify if it is a known contact. If it is a known contact, the server forwards the call directly to the user.

[0125] If the call is from an unknown contact, the server uses AI to inquire about the caller's purpose. The AI ​​performs speech-to-text conversion, natural language processing, and analyzes and classifies the request. During this process, an emotion engine extracts and analyzes emotional data from the transcribed information.

[0126] The emotion engine evaluates the caller's emotions, allowing the server to estimate whether the caller is angry, calm, or in an emergency. This enables more refined call forwarding decisions based on emotional information.

[0127] For example, if the server, through its emotion engine, determines that the caller is clearly feeling anxious, it is likely an urgent call and will prioritize transferring the call to the user. Similarly, even if the call is for business purposes, if the caller displays calm emotions, the system will automatically decline the call as usual, but if the caller is angry, it will respond accordingly.

[0128] Another feature of this system is its ability to record calls as needed and notify the caller that the call is being recorded. The recorded data is securely stored on the server for later review of the conversation with the caller.

[0129] By implementing this system, users will be able to handle calls in a way that is tailored to the caller's emotions and needs, resulting in a significant improvement in operational efficiency and the provision of a higher level of customer service.

[0130] The following describes the processing flow.

[0131] Step 1:

[0132] The server receives a call from a communication terminal. It checks the caller's phone number against a database and identifies the caller based on that information.

[0133] Step 2:

[0134] The server determines whether the caller is known or not. If the caller is known, the server forwards the call to the user's device.

[0135] Step 3:

[0136] If the server identifies the call as coming from an unknown caller, the AI ​​will verify the caller's request. The AI ​​will receive the caller's voice and convert it into text using speech recognition technology.

[0137] Step 4:

[0138] The server uses natural language processing technology to analyze the transcribed information and classify the message into categories such as sales, urgent, and other. It also uses an emotion engine to estimate the sender's emotional state.

[0139] Step 5:

[0140] The server determines whether to transfer the call to a user or handle it with an automated response, based on the caller's request and emotional state. If it is determined to be urgent, the call will be transferred to the user.

[0141] Step 6:

[0142] If the server determines it is a sales call, it will use an automated response system to send a rejection message to the caller. If the caller's mood is identified as negative, a special response will be provided.

[0143] Step 7:

[0144] If the user has enabled the recording function, the server will record the call and notify the caller. The recording data is securely stored on the server and can be reviewed as needed.

[0145] (Example 2)

[0146] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0147] In today's communication environment, recipients often receive a vast number of calls, making it difficult to quickly identify their content and urgency. This is especially true for calls from unknown callers, where judging their true intentions, emotions, and urgency is challenging, increasing the risk of missing important calls. Furthermore, there is a need for mechanisms to efficiently handle inappropriate sales calls and fraudulent activities. In addition, there is a demand for solutions that address these challenges while simultaneously improving user work efficiency and the quality of customer service.

[0148] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0149] In this invention, the server includes means for receiving signals from a communication device, means for acquiring and identifying caller information from an information aggregation means, voice analysis means for converting speech into text, natural language processing means for analyzing the transcribed information and classifying requirements, means for deciding whether or not to transfer the call to a user based on the caller's requirements and emotions, emotion analysis means for evaluating the caller's emotional state, and means for providing an automatic response to the caller. This allows for prioritizing important calls based on the caller's emotions and requirements and providing appropriate automatic responses, thereby improving the user's work efficiency and enabling higher quality customer service. Furthermore, by automatically classifying and addressing inappropriate calls, problematic calls can be processed efficiently.

[0150] "Communication equipment" refers to devices and infrastructure used to send and receive voice and data signals.

[0151] "Signal" refers to a means of representing information, including voice and data, that is transmitted or received through a communication device.

[0152] "Caller information" refers to data related to the number or identifier used to identify the sender during a phone call.

[0153] "Information storage means" refers to systems and devices that store data and information in a format that can be referenced as needed.

[0154] "Speech analysis means" refers to technologies and devices for converting speech data into text data.

[0155] "Transcripted information" refers to text data converted from speech, and means information used for information processing.

[0156] "Natural language processing methods" refer to technologies and systems that analyze human-understandable language and comprehend and classify its content.

[0157] "Emotional analysis methods" refer to technologies for extracting and evaluating the emotional state of a sender from their messages or audio data.

[0158] "User" refers to the person or organization that receives and processes calls using this system.

[0159] "Automatic response" refers to a process in which the system automatically responds based on predefined conditions.

[0160] To implement this invention, a system connected to a communication network is required, in which a server, terminal, and user each fulfill their respective roles.

[0161] First, the server receives signals from the communication device. A general-purpose communication protocol is used as the communication technology. The server retrieves information about the sender of the received signal from an information aggregation system and uses a database for identification.

[0162] Next, the server converts the speech into text using speech analysis tools. Specifically, technologies such as "speech recognition software" can be used. This transcribed information is then analyzed by natural language processing tools, and the requirements are classified.

[0163] The server then uses emotion analysis tools to evaluate the sender's emotional state. This process uses software tools such as an "emotion analysis engine" to extract and analyze the sender's emotions.

[0164] This system determines whether to transfer a call to a user based on analyzed requirements and sentiment information. Users can respond quickly to important calls. An automated response is provided to the caller as needed.

[0165] As a concrete example, consider a scenario where the user is a customer service representative. The user can receive notifications from the server regarding urgent customer inquiries and respond quickly. If the caller is determined to be potentially involved in inappropriate sales or fraud, an automated response will take appropriate action.

[0166] An example of a specific prompt for a generative AI model is, "Analyze the caller's voice and estimate their emotional state. Based on that information, select the appropriate call response." Such prompts allow the AI ​​to quickly analyze the caller's requirements and emotions, enabling the system to take the most appropriate action.

[0167] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0168] Step 1:

[0169] The server receives signals from the communication device. The input includes metadata such as the caller's number and identifier, and the timing of the call's initiation, obtained via the communication protocol. Based on this data, the server collects basic information about the caller and prepares for the next step. The output of this step is a dataset containing the caller's number and call metadata.

[0170] Step 2:

[0171] The server retrieves the caller's number from the information aggregation system and identifies whether it is a known or unknown contact. The input consists of known contact information registered in the database and the caller's number obtained in step 1. By matching it with the database, the server determines whether the caller is known. This ensures that calls from known callers are smoothly forwarded to the user. The output is tagged caller information as the identification result.

[0172] Step 3:

[0173] In the case of an unknown caller, the server uses speech analysis to convert the call content into text. As input, the server receives call audio data. This involves the specific operation of converting the audio into text data using speech recognition technology. This allows the server to obtain the call content in text format. The output is the transcribed call content.

[0174] Step 4:

[0175] The server uses a generative AI model to analyze information transcribed into text by natural language processing. The input is the text data obtained in step 3. Through prompts, the AI ​​model analyzes and classifies the text to identify requirements. As a result, information important to the user is extracted. The output is the classified requirements information.

[0176] Step 5:

[0177] The server evaluates the sender's emotions using emotion analysis tools. The text data obtained in step 4 is used as input. The emotion analysis engine analyzes the sender's emotional state and evaluates emotions such as anger and anxiety. The output is the estimated emotional state.

[0178] Step 6:

[0179] The server decides whether to transfer the call to the user based on the analyzed requirements and emotions. The input is the requirements information and emotion state obtained in steps 4 and 5. Based on this information, a decision is made on whether to transfer the call to the user or handle it with an automated response. The output is the result of the transfer method decision.

[0180] Step 7:

[0181] If necessary, the server implements a function to record the call and automatically notify the caller. The input is the call information for which the call transfer has been decided. Simultaneously with the start of recording, a message informing the caller that recording has begun is sent. The output of this step is the recorded call data and notification history.

[0182] (Application Example 2)

[0183] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0184] Conventional call management systems failed to adequately consider the caller's emotions or the urgency of the call, potentially leading to the inability to properly handle urgent calls. Furthermore, setting response policies based on the caller's emotional state was difficult, resulting in many calls being ineffectively processed. This highlighted the system's limitations, particularly in security-related situations where rapid and accurate responses were required.

[0185] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0186] In this invention, the server includes means for receiving calls from communication devices, means for obtaining and identifying caller characteristics from a database, means for converting speech to text, means for analyzing the textual information and classifying the requests, means for analyzing the caller's emotions and deciding whether or not to transfer the call to a user based on the caller's state, means for providing an automatic response to the caller, and means for evaluating the urgency and danger based on the emotion analysis results and setting call priorities. This makes it possible to set a response policy according to the caller's emotional state and to quickly determine the priority of calls according to their urgency.

[0187] A "communication device" is an electronic device used for making phone calls and data communications, and has the function of receiving and processing calls from users.

[0188] "Caller characteristics" refer to identifiable information associated with the individual or organization that initiated the call, including past call history and profile information recorded in the database.

[0189] "Voice conversion means" refers to a mechanism that converts call audio into text data, and has a process that converts voice signals into text information using automatic speech recognition technology.

[0190] "Natural language processing means" refers to a technology for analyzing textual information and classifying it according to specific requirements, and it has a process of understanding and analyzing natural language using artificial intelligence technology.

[0191] "Emotion" refers to the state of the speaker's feelings and sensations, and is usually analyzed as information extracted from audio or written text.

[0192] "Automatic response" is a function that automatically responds to the caller based on pre-set content and conditions, and is used to quickly respond to the needs of the person on the other end of the call.

[0193] "Means of setting priorities" refers to a mechanism that determines the processing priority and routes calls appropriately according to their importance and urgency, and includes control functions to ensure the system operates efficiently and effectively.

[0194] This invention relates to an advanced call management system that uses a communication device to receive calls from callers, analyze their content, and process them. When the server receives a call from a communication terminal, it retrieves the caller's characteristics from a database and identifies whether or not it is a known contact. If the call is from a known contact, the call is automatically forwarded to the user. If it is an unknown contact, the voice is converted to text using a voice conversion means. The converted text is analyzed by a natural language processing means to extract and classify the caller's request and emotions. Based on the emotions, the server sets call priorities and promptly notifies the user of high-priority calls.

[0195] The system also includes an automated call response function that, based on pre-set conditions, provides a response that matches the caller's emotions and requirements. This system incorporates various processes to ensure smooth call management, and these processes are executed by software programs on the server. Specifically, virtual libraries such as the "speech_recognition" library are used for speech conversion, and "EmotionEngine" is used for emotion analysis. Functions such as "CallRouter" are used for appropriate call routing.

[0196] Such systems are particularly useful in security services, enabling rapid response to users by quickly processing urgent calls. For example, in security services, they can quickly classify and notify users of calls indicating security threats, allowing security personnel to respond immediately.

[0197] The following is an example of a prompt message generated using a generative AI model:

[0198] Please provide example Python code for building an app that prioritizes emergency calls by converting call audio to text and identifying the caller's emotions. This should utilize speech recognition and sentiment analysis, and include emotion-based call routing functionality.

[0199] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0200] Step 1:

[0201] The server receives a call from a communication terminal. The input is voice data transmitted from the communication terminal. The server receives this voice data, compares the caller's number information with the database, and performs a calculation to determine whether or not it is a known contact. The output is the result of the determination of whether or not the caller is known.

[0202] Step 2:

[0203] The server uses a speech conversion method to convert the audio data into text data. Here, the input is the audio data from the previous step. Specifically, the "speech_recognition" library is used to convert the audio signal into text information. The output is the transcribed conversation content.

[0204] Step 3:

[0205] The server analyzes the transcribed information using natural language processing (NLP) to extract and classify the sender's message and emotions. The input is the text data obtained in step 2, and sentiment analysis is performed using "EmotionEngine". Data processing involves extracting emotional data from the text content and classifying the message based on that data. The output of this step is the classified message and the results of the sentiment analysis.

[0206] Step 4:

[0207] The server decides whether or not to transfer the call to the user based on the classified request and sentiment analysis results. The input is the request and sentiment data obtained in step 3. Specifically, calls are routed using logic to set priorities. "CallRouter" is used to process high-priority calls so that users are notified quickly. The output is the decision on whether or not to transfer the call.

[0208] Step 5:

[0209] The server will automatically respond to the caller as needed. The input is the call decision information from step 4. The server will respond to the caller based on a pre-configured automated response message. The output is the response sent to the caller.

[0210] Step 6:

[0211] The server prioritizes calls based on sentiment analysis results and records calls as needed. The input is the sentiment data from step 3. The recorded data is securely stored for later review. The output is the set priorities and recorded call data.

[0212] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0213] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0214] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0215] [Second Embodiment]

[0216] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0217] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0218] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0219] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0220] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0221] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0222] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0223] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0224] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0225] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0226] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0227] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0228] This invention is an AI-powered call management system designed to efficiently receive only the calls that the user needs. When the server receives a call from a communication terminal, it retrieves caller information from a database and identifies whether the caller is an acquaintance or not. If the caller is not an acquaintance, the server, through AI, confirms the requirements with the caller.

[0229] The AI ​​uses speech recognition technology to convert speech information into text. The converted text is then analyzed using natural language processing and classified as a sales call, a scam call, or for other purposes. Based on this classification information, the server decides whether to transfer the call to the user.

[0230] For example, if the server determines that a call is a sales call, it will use an automated response function to send a rejection message to the caller. On the other hand, if the caller is an acquaintance and the server recognizes that the matter is important, the server will transfer the call directly to the user's terminal.

[0231] Furthermore, users can optionally enable a recording function, which notifies the caller that the call is being recorded. The recording data is stored on the server for later reference.

[0232] This system is useful not only for personal use but also as a business telephone system for companies. By automatically processing unnecessary calls, it can improve work efficiency and enhance security.

[0233] The following describes the processing flow.

[0234] Step 1:

[0235] The server receives a call from a communication terminal. It detects the caller's phone number and retrieves the caller's information by searching the database based on that number.

[0236] Step 2:

[0237] Based on the caller information obtained by the server, the system checks if the caller is someone the user knows. If they are, the call is forwarded to the user's device.

[0238] Step 3:

[0239] The server connects the call to the AI ​​and asks the caller about their purpose. The AI ​​receives the caller's voice and converts it into text using speech recognition technology.

[0240] Step 4:

[0241] The server receives the message information transcribed into text from the AI ​​and performs natural language processing to analyze the message. This analysis classifies the message as sales, fraud, or other.

[0242] Step 5:

[0243] The server determines whether to transfer the call to the user based on the classification results. If it is determined to be an important call, it is transferred to the user's terminal so that the user can receive the call.

[0244] Step 6:

[0245] If the server determines that a call may be a sales call or a scam, it will use an automated response function to reject the caller with a pre-set message. If the recording option is enabled, it will notify the caller and record and save the conversation.

[0246] (Example 1)

[0247] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0248] In modern society, unnecessary and nuisance calls are frequent occurrences, often hindering the efficient operation of businesses and individuals. This can lead to missed important communications and the waste of time and resources. Therefore, there is a need for systems that efficiently receive only necessary calls and appropriately handle unnecessary ones.

[0249] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0250] In this invention, the server includes means for receiving calls from a communication device, means for acquiring and identifying caller information from an information storage device, and means for speech recognition that converts speech into text. This makes it possible for users to efficiently receive only important calls without being bothered by unnecessary calls.

[0251] "Communication equipment" is a general term for hardware or software used to send and receive data.

[0252] "Caller information" refers to the identification information of the caller used when making a phone call, and usually includes a phone number.

[0253] An "information storage device" is a database or storage system that stores data and allows it to be freely accessed as needed.

[0254] "Speech recognition means" is a general term for technologies or systems that convert speech input into string-formatted data.

[0255] "Stringified information" refers to text data converted by speech recognition technology.

[0256] "Natural language processing means" is a general term for systems that analyze string-based information to understand or classify its content.

[0257] "Recording function" refers to the process or function for saving call content as digital data.

[0258] "Means of automated response" refers to a system or process that mechanically responds to calls using pre-set messages.

[0259] This invention is a system that uses AI technology to streamline communication management. It primarily involves the coordinated operation of server, terminal, and user elements.

[0260] server

[0261] The server receives calls from communication devices and retrieves caller information from an information storage device. The server uses a database system to identify the caller information. Relational databases like MySQL are one option. Based on the received call information, the server processes the voice data.

[0262] Speech recognition means

[0263] The server uses speech recognition to convert the incoming speech into text. This can be done using commercial speech recognition software or cloud-based APIs. For example, Google Cloud Speech-to-Text can fulfill this role.

[0264] Natural language processing means

[0265] The converted strings are analyzed by a natural language processing system. Here, a generative AI model is used to classify requirements based on the text data. By using the OpenAI GPT model, it is possible to identify potential sales calls and scams with high accuracy.

[0266] Automatic response and recording function

[0267] Depending on the caller's requirements, the server will automatically respond. If it is determined to be a sales call, the server will automatically send a rejection message. Furthermore, users can enable the call recording function and save the recording to the server for later review.

[0268] Specific example

[0269] Suppose a user is in a coffee shop when they receive a call from an unknown number. In this case, the server quickly identifies it as a sales call and uses an automated response system to send a polite refusal message to the caller.

[0270] Example of a prompt

[0271] An example of a prompt for a generating AI model is: "When a call is received from an unknown caller, explain how to classify the call and notify the user. Also, specify how to handle sales calls."

[0272] This system offers users the advantage of receiving only necessary calls efficiently and securely, and also makes recording and managing calls easier.

[0273] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0274] Step 1:

[0275] The server receives a call from a communication device. At this time, it receives the call metadata (caller ID, time, date, etc.) as input and prepares to process the call. As output, the call information is recorded in the internal system and made available as input data for the next step.

[0276] Step 2:

[0277] The server retrieves and identifies source information from the information storage device. It uses the source number as input and accesses the database via an SQL query. Specifically, the server searches for records matching the source number, and the output is the detailed information of the corresponding record.

[0278] Step 3:

[0279] The server identifies whether the sender is an acquaintance. Based on information retrieved from the database, it compares it with the user's contact list and uses the sender information as input. If the sender is not an acquaintance, it outputs "Unknown sender" and proceeds to the next step.

[0280] Step 4:

[0281] For an unknown caller, the server uses voice recognition means to confirm the requirements. In this process, an automated voice message is played to the caller, and as a specific operation, a message saying "Please state your matter" is sent. The input is the voice of the caller, which is converted into text by voice recognition technology. As output, the requirements of the caller in string form are obtained.

[0282] Step 5:

[0283] The server analyzes the stringified information by natural language processing means and classifies the requirements. The input is text data, which is processed by a generative AI model. The specific operation is to classify the text data into business calls, fraud calls, and others. The output is the classification result, which is passed on to the next process.

[0284] Step 6:

[0285] Based on the classification result, the server decides whether to transfer the call to the user or reject it with an automated response. The classification result is used as input. Specifically, if it is determined to be a business call, a message saying "I am currently unable to connect" is sent to the caller. If it is identified as an important requirement, the output is to transfer the call to the user.

[0286] Step 7:

[0287] If the user wishes to record the call, the server enables the recording function. When the recording starts, the caller is notified that "This call is being recorded". As a specific operation, the recorded data is saved as output on the server and can be accessed by the user later.

[0288] (Application Example 1)

[0289] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0290] In today's communication environment, many unnecessary calls hinder the productivity of individuals and businesses. This is particularly true for customer support on e-commerce sites, where the wide range of customer inquiries makes it difficult to efficiently handle important calls. This can lead to delays in responding to critical inquiries that require a quick response.

[0291] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0292] In this invention, the server includes means for receiving calls from communication devices, means for obtaining and identifying caller information from a storage device, and means for speech recognition for converting speech into text. This makes it possible to automatically classify customer inquiries and quickly and efficiently transfer important inquiries to human operators.

[0293] "Communication equipment" refers to electronic devices used to send and receive voice and data over long distances.

[0294] A "call" is a request for communication made to another device or user using a communication device.

[0295] "Sender information" refers to data that includes the identifier and related information of the party that initiated the communication.

[0296] A "storage device" is a device that stores information and allows it to be reused as needed.

[0297] "Identification" is the process of recognizing a specific object and determining what it is.

[0298] "Speech recognition means" refers to technology that analyzes speech signals and converts them into text format.

[0299] A "text" is a composition in which meaningful words are arranged in a certain order.

[0300] The "natural language processing means" is a technology that analyzes texturized language data and understands its content and meaning.

[0301] "Classification" is the operation of grouping information according to specific criteria.

[0302] "Purpose" is a matter set as the aim of an action or process.

[0303] "User" refers to a person who uses or operates a system or device.

[0304] "Automatic response" refers to the reaction or response autonomously made by a machine or software.

[0305] "Operator" refers to a human being for operating a system or equipment. <00,00963>

[0306] In the system for realizing this invention, the server first receives a call from a communication device and plays the role of acquiring and identifying the caller information from the storage device. As hardware, a smartphone or a dedicated terminal as a communication device is used, and as software, the speech_recognition library of Python is introduced as speech recognition means. Then, the received voice is converted into text. The converted text is analyzed using a natural language processing library such as TextBlob and classified according to the purpose.

[0307] From the result of the analysis, the server identifies important inquiries and transfers them promptly to a human operator. Thereby, the efficiency of customer support on the e-commerce site can be greatly improved. As a specific example, when there is an inquiry from a customer about the shipment of a product, if it is determined to be important, it is transferred to the operator. At that time, a prompt sentence such as "Convert the voice data sent to customer support into text, classify the content of the inquiry, and perform corresponding processing." is used using a generative AI model.

[0308] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0309] Step 1:

[0310] The server receives a call from a communication device. At this point, the input is a call signal generated by the communication device, and the server takes this signal and outputs it as call information. Specifically, the server receives call information from the network using a communication protocol.

[0311] Step 2:

[0312] The server retrieves caller information from its storage device and identifies it. At this point, the input is the previously outputted call information, and a database query is generated based on the caller ID within it. This retrieves caller-related information from the database. This information is then provided as output. In its specific operation, SQL queries are used to retrieve the necessary data from the database.

[0313] Step 3:

[0314] The server uses speech recognition to convert received audio into text. The input is audio data, and by applying a speech recognition algorithm to it, the output is generated as text data. Specifically, the Python `speech_recognition` library is used for this process.

[0315] Step 4:

[0316] The server uses natural language processing to analyze the converted text and classify its purpose. The input is the text data generated in the previous step, and the output is the analyzed information classified by purpose. This process uses the TextBlob library for language analysis and classification.

[0317] Step 5:

[0318] The server determines, based on the classified information, whether to forward important inquiries to a human operator or to provide an automated response. The input is the classified information, and the output is the forwarding or response action. Specifically, important inquiries are forwarded to the operator's terminal using the API, while other inquiries generate automated response messages.

[0319] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0320] This invention is an advanced call management system using AI that optimizes calls received by users by analyzing various information, including the caller's emotions. When the server receives a call from a communication terminal, it compares the caller's number information with a database to identify if it is a known contact. If it is a known contact, the server forwards the call directly to the user.

[0321] If the call is from an unknown contact, the server uses AI to inquire about the caller's purpose. The AI ​​performs speech-to-text conversion, natural language processing, and analyzes and classifies the request. During this process, an emotion engine extracts and analyzes emotional data from the transcribed information.

[0322] The emotion engine evaluates the caller's emotions, allowing the server to estimate whether the caller is angry, calm, or in an emergency. This enables more refined call forwarding decisions based on emotional information.

[0323] For example, if the server, through its emotion engine, determines that the caller is clearly feeling anxious, it is likely an urgent call and will prioritize transferring the call to the user. Similarly, even if the call is for business purposes, if the caller displays calm emotions, the system will automatically decline the call as usual, but if the caller is angry, it will respond accordingly.

[0324] Another feature of this system is its ability to record calls as needed and notify the caller that the call is being recorded. The recorded data is securely stored on the server for later review of the conversation with the caller.

[0325] By implementing this system, users will be able to handle calls in a way that is tailored to the caller's emotions and needs, resulting in a significant improvement in operational efficiency and the provision of a higher level of customer service.

[0326] The following describes the processing flow.

[0327] Step 1:

[0328] The server receives a call from a communication terminal. It checks the caller's phone number against a database and identifies the caller based on that information.

[0329] Step 2:

[0330] The server determines whether the caller is known or not. If the caller is known, the server forwards the call to the user's device.

[0331] Step 3:

[0332] If the server identifies the call as coming from an unknown caller, the AI ​​will verify the caller's request. The AI ​​will receive the caller's voice and convert it into text using speech recognition technology.

[0333] Step 4:

[0334] The server uses natural language processing technology to analyze the transcribed information and classify the message into categories such as sales, urgent, and other. It also uses an emotion engine to estimate the sender's emotional state.

[0335] Step 5:

[0336] The server determines whether to transfer the call to a user or handle it with an automated response, based on the caller's request and emotional state. If it is determined to be urgent, the call will be transferred to the user.

[0337] Step 6:

[0338] If the server determines it is a sales call, it will use an automated response system to send a rejection message to the caller. If the caller's mood is identified as negative, a special response will be provided.

[0339] Step 7:

[0340] If the user has enabled the recording function, the server will record the call and notify the caller. The recording data is securely stored on the server and can be reviewed as needed.

[0341] (Example 2)

[0342] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0343] In today's communication environment, recipients often receive a vast number of calls, making it difficult to quickly identify their content and urgency. This is especially true for calls from unknown callers, where judging their true intentions, emotions, and urgency is challenging, increasing the risk of missing important calls. Furthermore, there is a need for mechanisms to efficiently handle inappropriate sales calls and fraudulent activities. In addition, there is a demand for solutions that address these challenges while simultaneously improving user work efficiency and the quality of customer service.

[0344] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0345] In this invention, the server includes means for receiving signals from a communication device, means for acquiring and identifying caller information from an information aggregation means, voice analysis means for converting speech into text, natural language processing means for analyzing the transcribed information and classifying requirements, means for deciding whether or not to transfer the call to a user based on the caller's requirements and emotions, emotion analysis means for evaluating the caller's emotional state, and means for providing an automatic response to the caller. This allows for prioritizing important calls based on the caller's emotions and requirements and providing appropriate automatic responses, thereby improving the user's work efficiency and enabling higher quality customer service. Furthermore, by automatically classifying and addressing inappropriate calls, problematic calls can be processed efficiently.

[0346] "Communication equipment" refers to devices and infrastructure used to send and receive voice and data signals.

[0347] "Signal" refers to a means of representing information, including voice and data, that is transmitted or received through a communication device.

[0348] "Caller information" refers to data related to the number or identifier used to identify the sender during a phone call.

[0349] "Information storage means" refers to systems and devices that store data and information in a format that can be referenced as needed.

[0350] "Speech analysis means" refers to technologies and devices for converting speech data into text data.

[0351] "Transcripted information" refers to text data converted from speech, and means information used for information processing.

[0352] "Natural language processing methods" refer to technologies and systems that analyze human-understandable language and comprehend and classify its content.

[0353] "Emotional analysis methods" refer to technologies for extracting and evaluating the emotional state of a sender from their messages or audio data.

[0354] "User" refers to the person or organization that receives and processes calls using this system.

[0355] "Automatic response" refers to a process in which the system automatically responds based on predefined conditions.

[0356] To implement this invention, a system connected to a communication network is required, in which a server, terminal, and user each fulfill their respective roles.

[0357] First, the server receives signals from the communication device. A general-purpose communication protocol is used as the communication technology. The server retrieves information about the sender of the received signal from an information aggregation system and uses a database for identification.

[0358] Next, the server converts the speech into text using speech analysis tools. Specifically, technologies such as "speech recognition software" can be used. This transcribed information is then analyzed by natural language processing tools, and the requirements are classified.

[0359] The server then uses emotion analysis tools to evaluate the sender's emotional state. This process uses software tools such as an "emotion analysis engine" to extract and analyze the sender's emotions.

[0360] This system determines whether to transfer a call to a user based on analyzed requirements and sentiment information. Users can respond quickly to important calls. An automated response is provided to the caller as needed.

[0361] As a concrete example, consider a scenario where the user is a customer service representative. The user can receive notifications from the server regarding urgent customer inquiries and respond quickly. If the caller is determined to be potentially involved in inappropriate sales or fraud, an automated response will take appropriate action.

[0362] An example of a specific prompt for a generative AI model is, "Analyze the caller's voice and estimate their emotional state. Based on that information, select the appropriate call response." Such prompts allow the AI ​​to quickly analyze the caller's requirements and emotions, enabling the system to take the most appropriate action.

[0363] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0364] Step 1:

[0365] The server receives signals from the communication device. The input includes metadata such as the caller's number and identifier, and the timing of the call's initiation, obtained via the communication protocol. Based on this data, the server collects basic information about the caller and prepares for the next step. The output of this step is a dataset containing the caller's number and call metadata.

[0366] Step 2:

[0367] The server retrieves the caller's number from the information aggregation system and identifies whether it is a known or unknown contact. The input consists of known contact information registered in the database and the caller's number obtained in step 1. By matching it with the database, the server determines whether the caller is known. This ensures that calls from known callers are smoothly forwarded to the user. The output is tagged caller information as the identification result.

[0368] Step 3:

[0369] In the case of an unknown caller, the server uses speech analysis to convert the call content into text. As input, the server receives call audio data. This involves the specific operation of converting the audio into text data using speech recognition technology. This allows the server to obtain the call content in text format. The output is the transcribed call content.

[0370] Step 4:

[0371] The server uses a generative AI model to analyze information transcribed into text by natural language processing. The input is the text data obtained in step 3. Through prompts, the AI ​​model analyzes and classifies the text to identify requirements. As a result, information important to the user is extracted. The output is the classified requirements information.

[0372] Step 5:

[0373] The server evaluates the sender's emotions using emotion analysis tools. The text data obtained in step 4 is used as input. The emotion analysis engine analyzes the sender's emotional state and evaluates emotions such as anger and anxiety. The output is the estimated emotional state.

[0374] Step 6:

[0375] The server decides whether to transfer the call to the user based on the analyzed requirements and emotions. The input is the requirements information and emotion state obtained in steps 4 and 5. Based on this information, a decision is made on whether to transfer the call to the user or handle it with an automated response. The output is the result of the transfer method decision.

[0376] Step 7:

[0377] If necessary, the server implements a function to record the call and automatically notify the caller. The input is the call information for which the call transfer has been decided. Simultaneously with the start of recording, a message informing the caller that recording has begun is sent. The output of this step is the recorded call data and notification history.

[0378] (Application Example 2)

[0379] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0380] Conventional call management systems failed to adequately consider the caller's emotions or the urgency of the call, potentially leading to the inability to properly handle urgent calls. Furthermore, setting response policies based on the caller's emotional state was difficult, resulting in many calls being ineffectively processed. This highlighted the system's limitations, particularly in security-related situations where rapid and accurate responses were required.

[0381] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0382] In this invention, the server includes means for receiving calls from communication devices, means for obtaining and identifying caller characteristics from a database, means for converting speech to text, means for analyzing the textual information and classifying the requests, means for analyzing the caller's emotions and deciding whether or not to transfer the call to a user based on the caller's state, means for providing an automatic response to the caller, and means for evaluating the urgency and danger based on the emotion analysis results and setting call priorities. This makes it possible to set a response policy according to the caller's emotional state and to quickly determine the priority of calls according to their urgency.

[0383] A "communication device" is an electronic device used for making phone calls and data communications, and has the function of receiving and processing calls from users.

[0384] "Caller characteristics" refer to identifiable information associated with the individual or organization that initiated the call, including past call history and profile information recorded in the database.

[0385] "Voice conversion means" refers to a mechanism that converts call audio into text data, and has a process that converts voice signals into text information using automatic speech recognition technology.

[0386] "Natural language processing means" refers to a technology for analyzing textual information and classifying it according to specific requirements, and it has a process of understanding and analyzing natural language using artificial intelligence technology.

[0387] "Emotion" refers to the state of the speaker's feelings and sensations, and is usually analyzed as information extracted from audio or written text.

[0388] "Automatic response" is a function that automatically responds to the caller based on pre-set content and conditions, and is used to quickly respond to the needs of the person on the other end of the call.

[0389] "Means of setting priorities" refers to a mechanism that determines the processing priority and routes calls appropriately according to their importance and urgency, and includes control functions to ensure the system operates efficiently and effectively.

[0390] This invention relates to an advanced call management system that uses a communication device to receive calls from callers, analyze their content, and process them. When the server receives a call from a communication terminal, it retrieves the caller's characteristics from a database and identifies whether or not it is a known contact. If the call is from a known contact, the call is automatically forwarded to the user. If it is an unknown contact, the voice is converted to text using a voice conversion means. The converted text is analyzed by a natural language processing means to extract and classify the caller's request and emotions. Based on the emotions, the server sets call priorities and promptly notifies the user of high-priority calls.

[0391] The system also includes an automated call response function that, based on pre-set conditions, provides a response that matches the caller's emotions and requirements. This system incorporates various processes to ensure smooth call management, and these processes are executed by software programs on the server. Specifically, virtual libraries such as the "speech_recognition" library are used for speech conversion, and "EmotionEngine" is used for emotion analysis. Functions such as "CallRouter" are used for appropriate call routing.

[0392] Such systems are particularly useful in security services, enabling rapid response to users by quickly processing urgent calls. For example, in security services, they can quickly classify and notify users of calls indicating security threats, allowing security personnel to respond immediately.

[0393] The following is an example of a prompt message generated using a generative AI model:

[0394] Please provide example Python code for building an app that prioritizes emergency calls by converting call audio to text and identifying the caller's emotions. This should utilize speech recognition and sentiment analysis, and include emotion-based call routing functionality.

[0395] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0396] Step 1:

[0397] The server receives a call from a communication terminal. The input is voice data transmitted from the communication terminal. The server receives this voice data, compares the caller's number information with the database, and performs a calculation to determine whether or not it is a known contact. The output is the result of the determination of whether or not the caller is known.

[0398] Step 2:

[0399] The server uses a speech conversion method to convert the audio data into text data. Here, the input is the audio data from the previous step. Specifically, the "speech_recognition" library is used to convert the audio signal into text information. The output is the transcribed conversation content.

[0400] Step 3:

[0401] The server analyzes the transcribed information using natural language processing (NLP) to extract and classify the sender's message and emotions. The input is the text data obtained in step 2, and sentiment analysis is performed using "EmotionEngine". Data processing involves extracting emotional data from the text content and classifying the message based on that data. The output of this step is the classified message and the results of the sentiment analysis.

[0402] Step 4:

[0403] The server decides whether or not to transfer the call to the user based on the classified request and sentiment analysis results. The input is the request and sentiment data obtained in step 3. Specifically, calls are routed using logic to set priorities. "CallRouter" is used to process high-priority calls so that users are notified quickly. The output is the decision on whether or not to transfer the call.

[0404] Step 5:

[0405] The server will automatically respond to the caller as needed. The input is the call decision information from step 4. The server will respond to the caller based on a pre-configured automated response message. The output is the response sent to the caller.

[0406] Step 6:

[0407] The server prioritizes calls based on sentiment analysis results and records calls as needed. The input is the sentiment data from step 3. The recorded data is securely stored for later review. The output is the set priorities and recorded call data.

[0408] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0409] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0410] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0411] [Third Embodiment]

[0412] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0413] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0414] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0415] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0416] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0417] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0418] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0419] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0420] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0421] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0422] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0423] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0424] This invention is an AI-powered call management system designed to efficiently receive only the calls that the user needs. When the server receives a call from a communication terminal, it retrieves caller information from a database and identifies whether the caller is an acquaintance or not. If the caller is not an acquaintance, the server, through AI, confirms the requirements with the caller.

[0425] The AI ​​uses speech recognition technology to convert speech information into text. The converted text is then analyzed using natural language processing and classified as a sales call, a scam call, or for other purposes. Based on this classification information, the server decides whether to transfer the call to the user.

[0426] For example, if the server determines that a call is a sales call, it will use an automated response function to send a rejection message to the caller. On the other hand, if the caller is an acquaintance and the server recognizes that the matter is important, the server will transfer the call directly to the user's terminal.

[0427] Furthermore, users can optionally enable a recording function, which notifies the caller that the call is being recorded. The recording data is stored on the server for later reference.

[0428] This system is useful not only for personal use but also as a business telephone system for companies. By automatically processing unnecessary calls, it can improve work efficiency and enhance security.

[0429] The following describes the processing flow.

[0430] Step 1:

[0431] The server receives a call from a communication terminal. It detects the caller's phone number and retrieves the caller's information by searching the database based on that number.

[0432] Step 2:

[0433] Based on the caller information obtained by the server, the system checks if the caller is someone the user knows. If they are, the call is forwarded to the user's device.

[0434] Step 3:

[0435] The server connects the call to the AI ​​and asks the caller about their purpose. The AI ​​receives the caller's voice and converts it into text using speech recognition technology.

[0436] Step 4:

[0437] The server receives the message information transcribed into text from the AI ​​and performs natural language processing to analyze the message. This analysis classifies the message as sales, fraud, or other.

[0438] Step 5:

[0439] The server determines whether to transfer the call to the user based on the classification results. If it is determined to be an important call, it is transferred to the user's terminal so that the user can receive the call.

[0440] Step 6:

[0441] If the server determines that a call may be a sales call or a scam, it will use an automated response function to reject the caller with a pre-set message. If the recording option is enabled, it will notify the caller and record and save the conversation.

[0442] (Example 1)

[0443] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0444] In modern society, unnecessary and nuisance calls are frequent occurrences, often hindering the efficient operation of businesses and individuals. This can lead to missed important communications and the waste of time and resources. Therefore, there is a need for systems that efficiently receive only necessary calls and appropriately handle unnecessary ones.

[0445] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0446] In this invention, the server includes means for receiving calls from a communication device, means for acquiring and identifying caller information from an information storage device, and means for speech recognition that converts speech into text. This makes it possible for users to efficiently receive only important calls without being bothered by unnecessary calls.

[0447] "Communication equipment" is a general term for hardware or software used to send and receive data.

[0448] "Caller information" refers to the identification information of the caller used when making a phone call, and usually includes a phone number.

[0449] An "information storage device" is a database or storage system that stores data and allows it to be freely accessed as needed.

[0450] "Speech recognition means" is a general term for technologies or systems that convert speech input into string-formatted data.

[0451] "Stringified information" refers to text data converted by speech recognition technology.

[0452] "Natural language processing means" is a general term for systems that analyze string-based information to understand or classify its content.

[0453] "Recording function" refers to the process or function for saving call content as digital data.

[0454] "Means of automated response" refers to a system or process that mechanically responds to calls using pre-set messages.

[0455] This invention is a system that uses AI technology to streamline communication management. It primarily involves the coordinated operation of server, terminal, and user elements.

[0456] server

[0457] The server receives calls from communication devices and retrieves caller information from an information storage device. The server uses a database system to identify the caller information. Relational databases like MySQL are one option. Based on the received call information, the server processes the voice data.

[0458] Speech recognition means

[0459] The server uses speech recognition to convert the incoming speech into text. This can be done using commercial speech recognition software or cloud-based APIs. For example, Google Cloud Speech-to-Text can fulfill this role.

[0460] Natural language processing means

[0461] The converted strings are analyzed by a natural language processing system. Here, a generative AI model is used to classify requirements based on the text data. By using the OpenAI GPT model, it is possible to identify potential sales calls and scams with high accuracy.

[0462] Automatic response and recording function

[0463] Depending on the caller's requirements, the server will automatically respond. If it is determined to be a sales call, the server will automatically send a rejection message. Furthermore, users can enable the call recording function and save the recording to the server for later review.

[0464] Specific example

[0465] Suppose a user is in a coffee shop when they receive a call from an unknown number. In this case, the server quickly identifies it as a sales call and uses an automated response system to send a polite refusal message to the caller.

[0466] Example of a prompt

[0467] An example of a prompt for a generating AI model is: "When a call is received from an unknown caller, explain how to classify the call and notify the user. Also, specify how to handle sales calls."

[0468] This system offers users the advantage of receiving only necessary calls efficiently and securely, and also makes recording and managing calls easier.

[0469] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0470] Step 1:

[0471] The server receives a call from a communication device. At this time, it receives the call metadata (caller ID, time, date, etc.) as input and prepares to process the call. As output, the call information is recorded in the internal system and made available as input data for the next step.

[0472] Step 2:

[0473] The server retrieves and identifies source information from the information storage device. It uses the source number as input and accesses the database via an SQL query. Specifically, the server searches for records matching the source number, and the output is the detailed information of the corresponding record.

[0474] Step 3:

[0475] The server identifies whether the sender is an acquaintance. Based on information retrieved from the database, it compares it with the user's contact list and uses the sender information as input. If the sender is not an acquaintance, it outputs "Unknown sender" and proceeds to the next step.

[0476] Step 4:

[0477] For callers who are not acquainted, the server uses speech recognition to confirm the requirements. In this process, an automated voice message is played to the caller, specifically the message, "Please state your request." The input is the caller's voice, which is converted into text using speech recognition technology. The output is the caller's requirements in text format.

[0478] Step 5:

[0479] The server analyzes the stringified information using natural language processing techniques and classifies the requirements. The input is text data, which is processed by a generative AI model. Specifically, it classifies the text data into sales calls, fraudulent calls, and other categories. The output, as the classification result, is then passed on to the next processing step.

[0480] Step 6:

[0481] Based on the classification results, the server decides whether to transfer the call to the user or to reject it with an automated response. The classification results are used as input; specifically, if it's determined to be a sales call, a message saying "We are currently unable to connect your call" is sent to the caller. If it's identified as an important matter, the call is transferred to the user as output.

[0482] Step 7:

[0483] If a user wishes to record a call, the server enables the recording function. When recording begins, the caller is notified that "this call is being recorded." Specifically, the recorded data is saved to the server as output and can be accessed by the user later.

[0484] (Application Example 1)

[0485] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0486] In today's communication environment, many unnecessary calls hinder the productivity of individuals and businesses. This is particularly true for customer support on e-commerce sites, where the wide range of customer inquiries makes it difficult to efficiently handle important calls. This can lead to delays in responding to critical inquiries that require a quick response.

[0487] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0488] In this invention, the server includes means for receiving calls from communication devices, means for obtaining and identifying caller information from a storage device, and means for speech recognition for converting speech into text. This makes it possible to automatically classify customer inquiries and quickly and efficiently transfer important inquiries to human operators.

[0489] "Communication equipment" refers to electronic devices used to send and receive voice and data over long distances.

[0490] A "call" is a request for communication made to another device or user using a communication device.

[0491] "Sender information" refers to data that includes the identifier and related information of the party that initiated the communication.

[0492] A "storage device" is a device that stores information and allows it to be reused as needed.

[0493] "Identification" is the process of recognizing a specific object and determining what it is.

[0494] "Speech recognition means" refers to technology that analyzes speech signals and converts them into text format.

[0495] A "text" is a composition in which meaningful words are arranged in a certain order.

[0496] "Natural language processing methods" are technologies that analyze text-based language data and understand its content and meaning.

[0497] "Classification" is the process of grouping information according to specific criteria.

[0498] "Purpose" refers to the goal set for an action or process.

[0499] A "user" is a person who uses or operates a system or device.

[0500] "Automated response" refers to a reaction or reply performed autonomously by a machine or software.

[0501] An "operator" refers to a person responsible for operating a system or piece of equipment.

[0502] In the system that implements this invention, the server is responsible for receiving calls from communication devices and identifying the caller by retrieving their information from storage. The hardware uses smartphones or dedicated terminals as communication devices, and the software uses the Python speech_recognition library as a speech recognition tool. The received speech is then converted into text. The converted text is analyzed using a natural language processing library such as TextBlob and classified according to its purpose.

[0503] Based on the analysis results, the server identifies important inquiries and quickly forwards them to human operators. This significantly improves the efficiency of customer support on e-commerce sites. For example, if a customer inquires about product shipping, it is deemed important and forwarded to an operator. At that time, a generative AI model is used to generate a prompt such as, "Convert the voice data received by customer support into text, classify the content of the inquiry, and take appropriate action."

[0504] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0505] Step 1:

[0506] The server receives a call from a communication device. At this point, the input is a call signal generated by the communication device, and the server takes this signal and outputs it as call information. Specifically, the server receives call information from the network using a communication protocol.

[0507] Step 2:

[0508] The server retrieves caller information from its storage device and identifies it. At this point, the input is the previously outputted call information, and a database query is generated based on the caller ID within it. This retrieves caller-related information from the database. This information is then provided as output. In its specific operation, SQL queries are used to retrieve the necessary data from the database.

[0509] Step 3:

[0510] The server uses speech recognition to convert received audio into text. The input is audio data, and by applying a speech recognition algorithm to it, the output is generated as text data. Specifically, the Python `speech_recognition` library is used for this process.

[0511] Step 4:

[0512] The server uses natural language processing to analyze the converted text and classify its purpose. The input is the text data generated in the previous step, and the output is the analyzed information classified by purpose. This process uses the TextBlob library for language analysis and classification.

[0513] Step 5:

[0514] The server determines, based on the classified information, whether to forward important inquiries to a human operator or to provide an automated response. The input is the classified information, and the output is the forwarding or response action. Specifically, important inquiries are forwarded to the operator's terminal using the API, while other inquiries generate automated response messages.

[0515] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0516] This invention is an advanced call management system using AI that optimizes calls received by users by analyzing various information, including the caller's emotions. When the server receives a call from a communication terminal, it compares the caller's number information with a database to identify if it is a known contact. If it is a known contact, the server forwards the call directly to the user.

[0517] If the call is from an unknown contact, the server uses AI to inquire about the caller's purpose. The AI ​​performs speech-to-text conversion, natural language processing, and analyzes and classifies the request. During this process, an emotion engine extracts and analyzes emotional data from the transcribed information.

[0518] The emotion engine evaluates the caller's emotions, allowing the server to estimate whether the caller is angry, calm, or in an emergency. This enables more refined call forwarding decisions based on emotional information.

[0519] For example, if the server, through its emotion engine, determines that the caller is clearly feeling anxious, it is likely an urgent call and will prioritize transferring the call to the user. Similarly, even if the call is for business purposes, if the caller displays calm emotions, the system will automatically decline the call as usual, but if the caller is angry, it will respond accordingly.

[0520] Another feature of this system is its ability to record calls as needed and notify the caller that the call is being recorded. The recorded data is securely stored on the server for later review of the conversation with the caller.

[0521] By implementing this system, users will be able to handle calls in a way that is tailored to the caller's emotions and needs, resulting in a significant improvement in operational efficiency and the provision of a higher level of customer service.

[0522] The following describes the processing flow.

[0523] Step 1:

[0524] The server receives a call from a communication terminal. It checks the caller's phone number against a database and identifies the caller based on that information.

[0525] Step 2:

[0526] The server determines whether the caller is known or not. If the caller is known, the server forwards the call to the user's device.

[0527] Step 3:

[0528] If the server identifies the call as coming from an unknown caller, the AI ​​will verify the caller's request. The AI ​​will receive the caller's voice and convert it into text using speech recognition technology.

[0529] Step 4:

[0530] The server uses natural language processing technology to analyze the transcribed information and classify the message into categories such as sales, urgent, and other. It also uses an emotion engine to estimate the sender's emotional state.

[0531] Step 5:

[0532] The server determines whether to transfer the call to a user or handle it with an automated response, based on the caller's request and emotional state. If it is determined to be urgent, the call will be transferred to the user.

[0533] Step 6:

[0534] If the server determines it is a sales call, it will use an automated response system to send a rejection message to the caller. If the caller's mood is identified as negative, a special response will be provided.

[0535] Step 7:

[0536] If the user has enabled the recording function, the server will record the call and notify the caller. The recording data is securely stored on the server and can be reviewed as needed.

[0537] (Example 2)

[0538] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0539] In today's communication environment, recipients often receive a vast number of calls, making it difficult to quickly identify their content and urgency. This is especially true for calls from unknown callers, where judging their true intentions, emotions, and urgency is challenging, increasing the risk of missing important calls. Furthermore, there is a need for mechanisms to efficiently handle inappropriate sales calls and fraudulent activities. In addition, there is a demand for solutions that address these challenges while simultaneously improving user work efficiency and the quality of customer service.

[0540] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0541] In this invention, the server includes means for receiving signals from a communication device, means for acquiring and identifying caller information from an information aggregation means, voice analysis means for converting speech into text, natural language processing means for analyzing the transcribed information and classifying requirements, means for deciding whether or not to transfer the call to a user based on the caller's requirements and emotions, emotion analysis means for evaluating the caller's emotional state, and means for providing an automatic response to the caller. This allows for prioritizing important calls based on the caller's emotions and requirements and providing appropriate automatic responses, thereby improving the user's work efficiency and enabling higher quality customer service. Furthermore, by automatically classifying and addressing inappropriate calls, problematic calls can be processed efficiently.

[0542] "Communication equipment" refers to devices and infrastructure used to send and receive voice and data signals.

[0543] "Signal" refers to a means of representing information, including voice and data, that is transmitted or received through a communication device.

[0544] "Caller information" refers to data related to the number or identifier used to identify the sender during a phone call.

[0545] "Information storage means" refers to systems and devices that store data and information in a format that can be referenced as needed.

[0546] "Speech analysis means" refers to technologies and devices for converting speech data into text data.

[0547] "Transcripted information" refers to text data converted from speech, and means information used for information processing.

[0548] "Natural language processing methods" refer to technologies and systems that analyze human-understandable language and comprehend and classify its content.

[0549] "Emotional analysis methods" refer to technologies for extracting and evaluating the emotional state of a sender from their messages or audio data.

[0550] "User" refers to the person or organization that receives and processes calls using this system.

[0551] "Automatic response" refers to a process in which the system automatically responds based on predefined conditions.

[0552] To implement this invention, a system connected to a communication network is required, in which a server, terminal, and user each fulfill their respective roles.

[0553] First, the server receives signals from the communication device. A general-purpose communication protocol is used as the communication technology. The server retrieves information about the sender of the received signal from an information aggregation system and uses a database for identification.

[0554] Next, the server converts the speech into text using speech analysis tools. Specifically, technologies such as "speech recognition software" can be used. This transcribed information is then analyzed by natural language processing tools, and the requirements are classified.

[0555] The server then uses emotion analysis tools to evaluate the sender's emotional state. This process uses software tools such as an "emotion analysis engine" to extract and analyze the sender's emotions.

[0556] This system determines whether to transfer a call to a user based on analyzed requirements and sentiment information. Users can respond quickly to important calls. An automated response is provided to the caller as needed.

[0557] As a concrete example, consider a scenario where the user is a customer service representative. The user can receive notifications from the server regarding urgent customer inquiries and respond quickly. If the caller is determined to be potentially involved in inappropriate sales or fraud, an automated response will take appropriate action.

[0558] An example of a specific prompt for a generative AI model is, "Analyze the caller's voice and estimate their emotional state. Based on that information, select the appropriate call response." Such prompts allow the AI ​​to quickly analyze the caller's requirements and emotions, enabling the system to take the most appropriate action.

[0559] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0560] Step 1:

[0561] The server receives signals from the communication device. The input includes metadata such as the caller's number and identifier, and the timing of the call's initiation, obtained via the communication protocol. Based on this data, the server collects basic information about the caller and prepares for the next step. The output of this step is a dataset containing the caller's number and call metadata.

[0562] Step 2:

[0563] The server retrieves the caller's number from the information aggregation system and identifies whether it is a known or unknown contact. The input consists of known contact information registered in the database and the caller's number obtained in step 1. By matching it with the database, the server determines whether the caller is known. This ensures that calls from known callers are smoothly forwarded to the user. The output is tagged caller information as the identification result.

[0564] Step 3:

[0565] In the case of an unknown caller, the server uses speech analysis to convert the call content into text. As input, the server receives call audio data. This involves the specific operation of converting the audio into text data using speech recognition technology. This allows the server to obtain the call content in text format. The output is the transcribed call content.

[0566] Step 4:

[0567] The server uses a generative AI model to analyze information transcribed into text by natural language processing. The input is the text data obtained in step 3. Through prompts, the AI ​​model analyzes and classifies the text to identify requirements. As a result, information important to the user is extracted. The output is the classified requirements information.

[0568] Step 5:

[0569] The server evaluates the sender's emotions using emotion analysis tools. The text data obtained in step 4 is used as input. The emotion analysis engine analyzes the sender's emotional state and evaluates emotions such as anger and anxiety. The output is the estimated emotional state.

[0570] Step 6:

[0571] The server decides whether to transfer the call to the user based on the analyzed requirements and emotions. The input is the requirements information and emotion state obtained in steps 4 and 5. Based on this information, a decision is made on whether to transfer the call to the user or handle it with an automated response. The output is the result of the transfer method decision.

[0572] Step 7:

[0573] If necessary, the server implements a function to record the call and automatically notify the caller. The input is the call information for which the call transfer has been decided. Simultaneously with the start of recording, a message informing the caller that recording has begun is sent. The output of this step is the recorded call data and notification history.

[0574] (Application Example 2)

[0575] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0576] Conventional call management systems failed to adequately consider the caller's emotions or the urgency of the call, potentially leading to the inability to properly handle urgent calls. Furthermore, setting response policies based on the caller's emotional state was difficult, resulting in many calls being ineffectively processed. This highlighted the system's limitations, particularly in security-related situations where rapid and accurate responses were required.

[0577] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0578] In this invention, the server includes means for receiving calls from communication devices, means for obtaining and identifying caller characteristics from a database, means for converting speech to text, means for analyzing the textual information and classifying the requests, means for analyzing the caller's emotions and deciding whether or not to transfer the call to a user based on the caller's state, means for providing an automatic response to the caller, and means for evaluating the urgency and danger based on the emotion analysis results and setting call priorities. This makes it possible to set a response policy according to the caller's emotional state and to quickly determine the priority of calls according to their urgency.

[0579] A "communication device" is an electronic device used for making phone calls and data communications, and has the function of receiving and processing calls from users.

[0580] "Caller characteristics" refer to identifiable information associated with the individual or organization that initiated the call, including past call history and profile information recorded in the database.

[0581] "Voice conversion means" refers to a mechanism that converts call audio into text data, and has a process that converts voice signals into text information using automatic speech recognition technology.

[0582] "Natural language processing means" refers to a technology for analyzing textual information and classifying it according to specific requirements, and it has a process of understanding and analyzing natural language using artificial intelligence technology.

[0583] "Emotion" refers to the state of the speaker's feelings and sensations, and is usually analyzed as information extracted from audio or written text.

[0584] "Automatic response" is a function that automatically responds to the caller based on pre-set content and conditions, and is used to quickly respond to the needs of the person on the other end of the call.

[0585] "Means of setting priorities" refers to a mechanism that determines the processing priority and routes calls appropriately according to their importance and urgency, and includes control functions to ensure the system operates efficiently and effectively.

[0586] This invention relates to an advanced call management system that uses a communication device to receive calls from callers, analyze their content, and process them. When the server receives a call from a communication terminal, it retrieves the caller's characteristics from a database and identifies whether or not it is a known contact. If the call is from a known contact, the call is automatically forwarded to the user. If it is an unknown contact, the voice is converted to text using a voice conversion means. The converted text is analyzed by a natural language processing means to extract and classify the caller's request and emotions. Based on the emotions, the server sets call priorities and promptly notifies the user of high-priority calls.

[0587] The system also includes an automated call response function that, based on pre-set conditions, provides a response that matches the caller's emotions and requirements. This system incorporates various processes to ensure smooth call management, and these processes are executed by software programs on the server. Specifically, virtual libraries such as the "speech_recognition" library are used for speech conversion, and "EmotionEngine" is used for emotion analysis. Functions such as "CallRouter" are used for appropriate call routing.

[0588] Such systems are particularly useful in security services, enabling rapid response to users by quickly processing urgent calls. For example, in security services, they can quickly classify and notify users of calls indicating security threats, allowing security personnel to respond immediately.

[0589] The following is an example of a prompt message generated using a generative AI model:

[0590] Please provide example Python code for building an app that prioritizes emergency calls by converting call audio to text and identifying the caller's emotions. This should utilize speech recognition and sentiment analysis, and include emotion-based call routing functionality.

[0591] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0592] Step 1:

[0593] The server receives a call from a communication terminal. The input is voice data transmitted from the communication terminal. The server receives this voice data, compares the caller's number information with the database, and performs a calculation to determine whether or not it is a known contact. The output is the result of the determination of whether or not the caller is known.

[0594] Step 2:

[0595] The server uses a speech conversion method to convert the audio data into text data. Here, the input is the audio data from the previous step. Specifically, the "speech_recognition" library is used to convert the audio signal into text information. The output is the transcribed conversation content.

[0596] Step 3:

[0597] The server analyzes the transcribed information using natural language processing (NLP) to extract and classify the sender's message and emotions. The input is the text data obtained in step 2, and sentiment analysis is performed using "EmotionEngine". Data processing involves extracting emotional data from the text content and classifying the message based on that data. The output of this step is the classified message and the results of the sentiment analysis.

[0598] Step 4:

[0599] The server decides whether or not to transfer the call to the user based on the classified request and sentiment analysis results. The input is the request and sentiment data obtained in step 3. Specifically, calls are routed using logic to set priorities. "CallRouter" is used to process high-priority calls so that users are notified quickly. The output is the decision on whether or not to transfer the call.

[0600] Step 5:

[0601] The server will automatically respond to the caller as needed. The input is the call decision information from step 4. The server will respond to the caller based on a pre-configured automated response message. The output is the response sent to the caller.

[0602] Step 6:

[0603] The server prioritizes calls based on sentiment analysis results and records calls as needed. The input is the sentiment data from step 3. The recorded data is securely stored for later review. The output is the set priorities and recorded call data.

[0604] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0605] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0606] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0607] [Fourth Embodiment]

[0608] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0609] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0610] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0611] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0612] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0613] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0614] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0615] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0616] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0617] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0618] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0619] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0620] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0621] This invention is an AI-powered call management system designed to efficiently receive only the calls that the user needs. When the server receives a call from a communication terminal, it retrieves caller information from a database and identifies whether the caller is an acquaintance or not. If the caller is not an acquaintance, the server, through AI, confirms the requirements with the caller.

[0622] The AI ​​uses speech recognition technology to convert speech information into text. The converted text is then analyzed using natural language processing and classified as a sales call, a scam call, or for other purposes. Based on this classification information, the server decides whether to transfer the call to the user.

[0623] For example, if the server determines that a call is a sales call, it will use an automated response function to send a rejection message to the caller. On the other hand, if the caller is an acquaintance and the server recognizes that the matter is important, the server will transfer the call directly to the user's terminal.

[0624] Furthermore, users can optionally enable a recording function, which notifies the caller that the call is being recorded. The recording data is stored on the server for later reference.

[0625] This system is useful not only for personal use but also as a business telephone system for companies. By automatically processing unnecessary calls, it can improve work efficiency and enhance security.

[0626] The following describes the processing flow.

[0627] Step 1:

[0628] The server receives a call from a communication terminal. It detects the caller's phone number and retrieves the caller's information by searching the database based on that number.

[0629] Step 2:

[0630] Based on the caller information obtained by the server, the system checks if the caller is someone the user knows. If they are, the call is forwarded to the user's device.

[0631] Step 3:

[0632] The server connects the call to the AI ​​and asks the caller about their purpose. The AI ​​receives the caller's voice and converts it into text using speech recognition technology.

[0633] Step 4:

[0634] The server receives the message information transcribed into text from the AI ​​and performs natural language processing to analyze the message. This analysis classifies the message as sales, fraud, or other.

[0635] Step 5:

[0636] The server determines whether to transfer the call to the user based on the classification results. If it is determined to be an important call, it is transferred to the user's terminal so that the user can receive the call.

[0637] Step 6:

[0638] If the server determines that a call may be a sales call or a scam, it will use an automated response function to reject the caller with a pre-set message. If the recording option is enabled, it will notify the caller and record and save the conversation.

[0639] (Example 1)

[0640] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0641] In modern society, unnecessary and nuisance calls are frequent occurrences, often hindering the efficient operation of businesses and individuals. This can lead to missed important communications and the waste of time and resources. Therefore, there is a need for systems that efficiently receive only necessary calls and appropriately handle unnecessary ones.

[0642] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0643] In this invention, the server includes means for receiving calls from a communication device, means for acquiring and identifying caller information from an information storage device, and means for speech recognition that converts speech into text. This makes it possible for users to efficiently receive only important calls without being bothered by unnecessary calls.

[0644] "Communication equipment" is a general term for hardware or software used to send and receive data.

[0645] "Caller information" refers to the identification information of the caller used when making a phone call, and usually includes a phone number.

[0646] An "information storage device" is a database or storage system that stores data and allows it to be freely accessed as needed.

[0647] "Speech recognition means" is a general term for technologies or systems that convert speech input into string-formatted data.

[0648] "Stringified information" refers to text data converted by speech recognition technology.

[0649] "Natural language processing means" is a general term for systems that analyze string-based information to understand or classify its content.

[0650] "Recording function" refers to the process or function for saving call content as digital data.

[0651] "Means of automated response" refers to a system or process that mechanically responds to calls using pre-set messages.

[0652] This invention is a system that uses AI technology to streamline communication management. It primarily involves the coordinated operation of server, terminal, and user elements.

[0653] server

[0654] The server receives calls from communication devices and retrieves caller information from an information storage device. The server uses a database system to identify the caller information. Relational databases like MySQL are one option. Based on the received call information, the server processes the voice data.

[0655] Speech recognition means

[0656] The server uses speech recognition to convert the incoming speech into text. This can be done using commercial speech recognition software or cloud-based APIs. For example, Google Cloud Speech-to-Text can fulfill this role.

[0657] Natural language processing means

[0658] The converted strings are analyzed by a natural language processing system. Here, a generative AI model is used to classify requirements based on the text data. By using the OpenAI GPT model, it is possible to identify potential sales calls and scams with high accuracy.

[0659] Automatic response and recording function

[0660] Depending on the caller's requirements, the server will automatically respond. If it is determined to be a sales call, the server will automatically send a rejection message. Furthermore, users can enable the call recording function and save the recording to the server for later review.

[0661] Specific example

[0662] Suppose a user is in a coffee shop when they receive a call from an unknown number. In this case, the server quickly identifies it as a sales call and uses an automated response system to send a polite refusal message to the caller.

[0663] Example of a prompt

[0664] An example of a prompt for a generating AI model is: "When a call is received from an unknown caller, explain how to classify the call and notify the user. Also, specify how to handle sales calls."

[0665] This system offers users the advantage of receiving only necessary calls efficiently and securely, and also makes recording and managing calls easier.

[0666] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0667] Step 1:

[0668] The server receives a call from a communication device. At this time, it receives the call metadata (caller ID, time, date, etc.) as input and prepares to process the call. As output, the call information is recorded in the internal system and made available as input data for the next step.

[0669] Step 2:

[0670] The server retrieves and identifies source information from the information storage device. It uses the source number as input and accesses the database via an SQL query. Specifically, the server searches for records matching the source number, and the output is the detailed information of the corresponding record.

[0671] Step 3:

[0672] The server identifies whether the sender is an acquaintance. Based on information retrieved from the database, it compares it with the user's contact list and uses the sender information as input. If the sender is not an acquaintance, it outputs "Unknown sender" and proceeds to the next step.

[0673] Step 4:

[0674] For callers who are not acquainted, the server uses speech recognition to confirm the requirements. In this process, an automated voice message is played to the caller, specifically the message, "Please state your request." The input is the caller's voice, which is converted into text using speech recognition technology. The output is the caller's requirements in text format.

[0675] Step 5:

[0676] The server analyzes the stringified information using natural language processing techniques and classifies the requirements. The input is text data, which is processed by a generative AI model. Specifically, it classifies the text data into sales calls, fraudulent calls, and other categories. The output, as the classification result, is then passed on to the next processing step.

[0677] Step 6:

[0678] Based on the classification results, the server decides whether to transfer the call to the user or to reject it with an automated response. The classification results are used as input; specifically, if it's determined to be a sales call, a message saying "We are currently unable to connect your call" is sent to the caller. If it's identified as an important matter, the call is transferred to the user as output.

[0679] Step 7:

[0680] If a user wishes to record a call, the server enables the recording function. When recording begins, the caller is notified that "this call is being recorded." Specifically, the recorded data is saved to the server as output and can be accessed by the user later.

[0681] (Application Example 1)

[0682] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0683] In today's communication environment, many unnecessary calls hinder the productivity of individuals and businesses. This is particularly true for customer support on e-commerce sites, where the wide range of customer inquiries makes it difficult to efficiently handle important calls. This can lead to delays in responding to critical inquiries that require a quick response.

[0684] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0685] In this invention, the server includes means for receiving calls from communication devices, means for obtaining and identifying caller information from a storage device, and means for speech recognition for converting speech into text. This makes it possible to automatically classify customer inquiries and quickly and efficiently transfer important inquiries to human operators.

[0686] "Communication equipment" refers to electronic devices used to send and receive voice and data over long distances.

[0687] A "call" is a request for communication made to another device or user using a communication device.

[0688] "Sender information" refers to data that includes the identifier and related information of the party that initiated the communication.

[0689] A "storage device" is a device that stores information and allows it to be reused as needed.

[0690] "Identification" is the process of recognizing a specific object and determining what it is.

[0691] "Speech recognition means" refers to technology that analyzes speech signals and converts them into text format.

[0692] A "text" is a composition in which meaningful words are arranged in a certain order.

[0693] "Natural language processing methods" are technologies that analyze text-based language data and understand its content and meaning.

[0694] "Classification" is the process of grouping information according to specific criteria.

[0695] "Purpose" refers to the goal set for an action or process.

[0696] A "user" is a person who uses or operates a system or device.

[0697] "Automated response" refers to a reaction or reply performed autonomously by a machine or software.

[0698] An "operator" refers to a person responsible for operating a system or piece of equipment.

[0699] In the system that implements this invention, the server is responsible for receiving calls from communication devices and identifying the caller by retrieving their information from storage. The hardware uses smartphones or dedicated terminals as communication devices, and the software uses the Python speech_recognition library as a speech recognition tool. The received speech is then converted into text. The converted text is analyzed using a natural language processing library such as TextBlob and classified according to its purpose.

[0700] Based on the analysis results, the server identifies important inquiries and quickly forwards them to human operators. This significantly improves the efficiency of customer support on e-commerce sites. For example, if a customer inquires about product shipping, it is deemed important and forwarded to an operator. At that time, a generative AI model is used to generate a prompt such as, "Convert the voice data received by customer support into text, classify the content of the inquiry, and take appropriate action."

[0701] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0702] Step 1:

[0703] The server receives a call from a communication device. At this point, the input is a call signal generated by the communication device, and the server takes this signal and outputs it as call information. Specifically, the server receives call information from the network using a communication protocol.

[0704] Step 2:

[0705] The server retrieves caller information from its storage device and identifies it. At this point, the input is the previously outputted call information, and a database query is generated based on the caller ID within it. This retrieves caller-related information from the database. This information is then provided as output. In its specific operation, SQL queries are used to retrieve the necessary data from the database.

[0706] Step 3:

[0707] The server uses speech recognition to convert received audio into text. The input is audio data, and by applying a speech recognition algorithm to it, the output is generated as text data. Specifically, the Python `speech_recognition` library is used for this process.

[0708] Step 4:

[0709] The server uses natural language processing to analyze the converted text and classify its purpose. The input is the text data generated in the previous step, and the output is the analyzed information classified by purpose. This process uses the TextBlob library for language analysis and classification.

[0710] Step 5:

[0711] The server determines, based on the classified information, whether to forward important inquiries to a human operator or to provide an automated response. The input is the classified information, and the output is the forwarding or response action. Specifically, important inquiries are forwarded to the operator's terminal using the API, while other inquiries generate automated response messages.

[0712] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0713] This invention is an advanced call management system using AI that optimizes calls received by users by analyzing various information, including the caller's emotions. When the server receives a call from a communication terminal, it compares the caller's number information with a database to identify if it is a known contact. If it is a known contact, the server forwards the call directly to the user.

[0714] If the call is from an unknown contact, the server uses AI to inquire about the caller's purpose. The AI ​​performs speech-to-text conversion, natural language processing, and analyzes and classifies the request. During this process, an emotion engine extracts and analyzes emotional data from the transcribed information.

[0715] The emotion engine evaluates the caller's emotions, allowing the server to estimate whether the caller is angry, calm, or in an emergency. This enables more refined call forwarding decisions based on emotional information.

[0716] For example, if the server, through its emotion engine, determines that the caller is clearly feeling anxious, it is likely an urgent call and will prioritize transferring the call to the user. Similarly, even if the call is for business purposes, if the caller displays calm emotions, the system will automatically decline the call as usual, but if the caller is angry, it will respond accordingly.

[0717] Another feature of this system is its ability to record calls as needed and notify the caller that the call is being recorded. The recorded data is securely stored on the server for later review of the conversation with the caller.

[0718] By implementing this system, users will be able to handle calls in a way that is tailored to the caller's emotions and needs, resulting in a significant improvement in operational efficiency and the provision of a higher level of customer service.

[0719] The following describes the processing flow.

[0720] Step 1:

[0721] The server receives a call from a communication terminal. It checks the caller's phone number against a database and identifies the caller based on that information.

[0722] Step 2:

[0723] The server determines whether the caller is known or not. If the caller is known, the server forwards the call to the user's device.

[0724] Step 3:

[0725] If the server identifies the call as coming from an unknown caller, the AI ​​will verify the caller's request. The AI ​​will receive the caller's voice and convert it into text using speech recognition technology.

[0726] Step 4:

[0727] The server uses natural language processing technology to analyze the transcribed information and classify the message into categories such as sales, urgent, and other. It also uses an emotion engine to estimate the sender's emotional state.

[0728] Step 5:

[0729] The server determines whether to transfer the call to a user or handle it with an automated response, based on the caller's request and emotional state. If it is determined to be urgent, the call will be transferred to the user.

[0730] Step 6:

[0731] If the server determines it is a sales call, it will use an automated response system to send a rejection message to the caller. If the caller's mood is identified as negative, a special response will be provided.

[0732] Step 7:

[0733] If the user has enabled the recording function, the server will record the call and notify the caller. The recording data is securely stored on the server and can be reviewed as needed.

[0734] (Example 2)

[0735] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0736] In today's communication environment, recipients often receive a vast number of calls, making it difficult to quickly identify their content and urgency. This is especially true for calls from unknown callers, where judging their true intentions, emotions, and urgency is challenging, increasing the risk of missing important calls. Furthermore, there is a need for mechanisms to efficiently handle inappropriate sales calls and fraudulent activities. In addition, there is a demand for solutions that address these challenges while simultaneously improving user work efficiency and the quality of customer service.

[0737] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0738] In this invention, the server includes means for receiving signals from a communication device, means for acquiring and identifying caller information from an information aggregation means, voice analysis means for converting speech into text, natural language processing means for analyzing the transcribed information and classifying requirements, means for deciding whether or not to transfer the call to a user based on the caller's requirements and emotions, emotion analysis means for evaluating the caller's emotional state, and means for providing an automatic response to the caller. This allows for prioritizing important calls based on the caller's emotions and requirements and providing appropriate automatic responses, thereby improving the user's work efficiency and enabling higher quality customer service. Furthermore, by automatically classifying and addressing inappropriate calls, problematic calls can be processed efficiently.

[0739] "Communication equipment" refers to devices and infrastructure used to send and receive voice and data signals.

[0740] "Signal" refers to a means of representing information, including voice and data, that is transmitted or received through a communication device.

[0741] "Caller information" refers to data related to the number or identifier used to identify the sender during a phone call.

[0742] "Information storage means" refers to systems and devices that store data and information in a format that can be referenced as needed.

[0743] "Speech analysis means" refers to technologies and devices for converting speech data into text data.

[0744] "Transcripted information" refers to text data converted from speech, and means information used for information processing.

[0745] "Natural language processing methods" refer to technologies and systems that analyze human-understandable language and comprehend and classify its content.

[0746] "Emotional analysis methods" refer to technologies for extracting and evaluating the emotional state of a sender from their messages or audio data.

[0747] "User" refers to the person or organization that receives and processes calls using this system.

[0748] "Automatic response" refers to a process in which the system automatically responds based on predefined conditions.

[0749] To implement this invention, a system connected to a communication network is required, in which a server, terminal, and user each fulfill their respective roles.

[0750] First, the server receives signals from the communication device. A general-purpose communication protocol is used as the communication technology. The server retrieves information about the sender of the received signal from an information aggregation system and uses a database for identification.

[0751] Next, the server converts the speech into text using speech analysis tools. Specifically, technologies such as "speech recognition software" can be used. This transcribed information is then analyzed by natural language processing tools, and the requirements are classified.

[0752] The server then uses emotion analysis tools to evaluate the sender's emotional state. This process uses software tools such as an "emotion analysis engine" to extract and analyze the sender's emotions.

[0753] This system determines whether to transfer a call to a user based on analyzed requirements and sentiment information. Users can respond quickly to important calls. An automated response is provided to the caller as needed.

[0754] As a concrete example, consider a scenario where the user is a customer service representative. The user can receive notifications from the server regarding urgent customer inquiries and respond quickly. If the caller is determined to be potentially involved in inappropriate sales or fraud, an automated response will take appropriate action.

[0755] An example of a specific prompt for a generative AI model is, "Analyze the caller's voice and estimate their emotional state. Based on that information, select the appropriate call response." Such prompts allow the AI ​​to quickly analyze the caller's requirements and emotions, enabling the system to take the most appropriate action.

[0756] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0757] Step 1:

[0758] The server receives signals from the communication device. The input includes metadata such as the caller's number and identifier, and the timing of the call's initiation, obtained via the communication protocol. Based on this data, the server collects basic information about the caller and prepares for the next step. The output of this step is a dataset containing the caller's number and call metadata.

[0759] Step 2:

[0760] The server retrieves the caller's number from the information aggregation system and identifies whether it is a known or unknown contact. The input consists of known contact information registered in the database and the caller's number obtained in step 1. By matching it with the database, the server determines whether the caller is known. This ensures that calls from known callers are smoothly forwarded to the user. The output is tagged caller information as the identification result.

[0761] Step 3:

[0762] In the case of an unknown caller, the server uses speech analysis to convert the call content into text. As input, the server receives call audio data. This involves the specific operation of converting the audio into text data using speech recognition technology. This allows the server to obtain the call content in text format. The output is the transcribed call content.

[0763] Step 4:

[0764] The server uses a generative AI model to analyze information transcribed into text by natural language processing. The input is the text data obtained in step 3. Through prompts, the AI ​​model analyzes and classifies the text to identify requirements. As a result, information important to the user is extracted. The output is the classified requirements information.

[0765] Step 5:

[0766] The server evaluates the sender's emotions using emotion analysis tools. The text data obtained in step 4 is used as input. The emotion analysis engine analyzes the sender's emotional state and evaluates emotions such as anger and anxiety. The output is the estimated emotional state.

[0767] Step 6:

[0768] The server decides whether to transfer the call to the user based on the analyzed requirements and emotions. The input is the requirements information and emotion state obtained in steps 4 and 5. Based on this information, a decision is made on whether to transfer the call to the user or handle it with an automated response. The output is the result of the transfer method decision.

[0769] Step 7:

[0770] If necessary, the server implements a function to record the call and automatically notify the caller. The input is the call information for which the call transfer has been decided. Simultaneously with the start of recording, a message informing the caller that recording has begun is sent. The output of this step is the recorded call data and notification history.

[0771] (Application Example 2)

[0772] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0773] Conventional call management systems failed to adequately consider the caller's emotions or the urgency of the call, potentially leading to the inability to properly handle urgent calls. Furthermore, setting response policies based on the caller's emotional state was difficult, resulting in many calls being ineffectively processed. This highlighted the system's limitations, particularly in security-related situations where rapid and accurate responses were required.

[0774] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0775] In this invention, the server includes means for receiving calls from communication devices, means for obtaining and identifying caller characteristics from a database, means for converting speech to text, means for analyzing the textual information and classifying the requests, means for analyzing the caller's emotions and deciding whether or not to transfer the call to a user based on the caller's state, means for providing an automatic response to the caller, and means for evaluating the urgency and danger based on the emotion analysis results and setting call priorities. This makes it possible to set a response policy according to the caller's emotional state and to quickly determine the priority of calls according to their urgency.

[0776] A "communication device" is an electronic device used for making phone calls and data communications, and has the function of receiving and processing calls from users.

[0777] "Caller characteristics" refer to identifiable information associated with the individual or organization that initiated the call, including past call history and profile information recorded in the database.

[0778] "Voice conversion means" refers to a mechanism that converts call audio into text data, and has a process that converts voice signals into text information using automatic speech recognition technology.

[0779] "Natural language processing means" refers to a technology for analyzing textual information and classifying it according to specific requirements, and it has a process of understanding and analyzing natural language using artificial intelligence technology.

[0780] "Emotion" refers to the state of the speaker's feelings and sensations, and is usually analyzed as information extracted from audio or written text.

[0781] "Automatic response" is a function that automatically responds to the caller based on pre-set content and conditions, and is used to quickly respond to the needs of the person on the other end of the call.

[0782] "Means of setting priorities" refers to a mechanism that determines the processing priority and routes calls appropriately according to their importance and urgency, and includes control functions to ensure the system operates efficiently and effectively.

[0783] This invention relates to an advanced call management system that uses a communication device to receive calls from callers, analyze their content, and process them. When the server receives a call from a communication terminal, it retrieves the caller's characteristics from a database and identifies whether or not it is a known contact. If the call is from a known contact, the call is automatically forwarded to the user. If it is an unknown contact, the voice is converted to text using a voice conversion means. The converted text is analyzed by a natural language processing means to extract and classify the caller's request and emotions. Based on the emotions, the server sets call priorities and promptly notifies the user of high-priority calls.

[0784] The system also includes an automated call response function that, based on pre-set conditions, provides a response that matches the caller's emotions and requirements. This system incorporates various processes to ensure smooth call management, and these processes are executed by software programs on the server. Specifically, virtual libraries such as the "speech_recognition" library are used for speech conversion, and "EmotionEngine" is used for emotion analysis. Functions such as "CallRouter" are used for appropriate call routing.

[0785] Such systems are particularly useful in security services, enabling rapid response to users by quickly processing urgent calls. For example, in security services, they can quickly classify and notify users of calls indicating security threats, allowing security personnel to respond immediately.

[0786] The following is an example of a prompt message generated using a generative AI model:

[0787] Please provide example Python code for building an app that prioritizes emergency calls by converting call audio to text and identifying the caller's emotions. This should utilize speech recognition and sentiment analysis, and include emotion-based call routing functionality.

[0788] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0789] Step 1:

[0790] The server receives a call from a communication terminal. The input is voice data transmitted from the communication terminal. The server receives this voice data, compares the caller's number information with the database, and performs a calculation to determine whether or not it is a known contact. The output is the result of the determination of whether or not the caller is known.

[0791] Step 2:

[0792] The server uses a speech conversion method to convert the audio data into text data. Here, the input is the audio data from the previous step. Specifically, the "speech_recognition" library is used to convert the audio signal into text information. The output is the transcribed conversation content.

[0793] Step 3:

[0794] The server analyzes the transcribed information using natural language processing (NLP) to extract and classify the sender's message and emotions. The input is the text data obtained in step 2, and sentiment analysis is performed using "EmotionEngine". Data processing involves extracting emotional data from the text content and classifying the message based on that data. The output of this step is the classified message and the results of the sentiment analysis.

[0795] Step 4:

[0796] The server decides whether or not to transfer the call to the user based on the classified request and sentiment analysis results. The input is the request and sentiment data obtained in step 3. Specifically, calls are routed using logic to set priorities. "CallRouter" is used to process high-priority calls so that users are notified quickly. The output is the decision on whether or not to transfer the call.

[0797] Step 5:

[0798] The server will automatically respond to the caller as needed. The input is the call decision information from step 4. The server will respond to the caller based on a pre-configured automated response message. The output is the response sent to the caller.

[0799] Step 6:

[0800] The server prioritizes calls based on sentiment analysis results and records calls as needed. The input is the sentiment data from step 3. The recorded data is securely stored for later review. The output is the set priorities and recorded call data.

[0801] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0802] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0803] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0804] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0805] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0806] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0807] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0808] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0809] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0810] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0811] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0812] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0813] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0814] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0815] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0816] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0817] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0818] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0819] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0820] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0821] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0822] The following is further disclosed regarding the embodiments described above.

[0823] (Claim 1)

[0824] A means of receiving calls from a communication terminal,

[0825] A means of obtaining and identifying sender information from a database,

[0826] A speech recognition method that converts speech to text,

[0827] A natural language processing tool that analyzes transcribed information and classifies requirements,

[0828] A means of determining whether or not to transfer a call to the user based on the caller's requirements,

[0829] A system that includes means for providing an automated response to the caller.

[0830] (Claim 2)

[0831] The system according to claim 1, further comprising means for recording a call and notifying the caller of that fact.

[0832] (Claim 3)

[0833] The system according to claim 1, comprising means for automatically rejecting calls if they are potentially business or fraudulent, based on information classified by natural language processing means.

[0834] "Example 1"

[0835] (Claim 1)

[0836] Means for receiving calls from communication devices,

[0837] A means for obtaining and identifying source information from an information storage device,

[0838] A speech recognition means for converting speech into text,

[0839] A natural language processing method that analyzes stringified information and classifies requirements,

[0840] A means of determining whether or not to transfer a call to the user based on the caller's requirements,

[0841] A means of automatically responding to the caller,

[0842] A means of enabling the recording function and notifying the sender of this fact,

[0843] A system that includes means to automatically reject calls based on analyzed information and specific requirements.

[0844] (Claim 2)

[0845] The system according to claim 1, further comprising means for recording a call and notifying the caller accordingly.

[0846] (Claim 3)

[0847] The system according to claim 1, comprising means for automatically rejecting a call when there are specific requirements based on information classified by natural language processing means.

[0848] "Application Example 1"

[0849] (Claim 1)

[0850] A means for receiving calls from communication devices,

[0851] A means for obtaining and identifying caller information from a storage device,

[0852] A speech recognition method that converts speech into text,

[0853] A natural language processing method that analyzes the converted information and classifies the objective,

[0854] A means of deciding whether or not to transfer a call to a user based on the caller's purpose,

[0855] A means of providing an automated response to the caller,

[0856] A system that includes means for quickly transferring important inquiries to human operators.

[0857] (Claim 2)

[0858] The system according to claim 1, further comprising means for recording a call and notifying the caller of that fact.

[0859] (Claim 3)

[0860] The system according to claim 1, comprising means for automatically rejecting calls that may be commercial solicitations or fraudulent based on information classified by natural language processing means.

[0861] "Example 2 of combining an emotion engine"

[0862] (Claim 1)

[0863] Means for receiving signals from a communication device,

[0864] A means for obtaining and identifying sender information from an information aggregation means,

[0865] A speech analysis means for converting speech into text,

[0866] A natural language processing means that analyzes transcribed information and classifies requirements,

[0867] A means of determining whether or not to transfer a call to a user based on the caller's requirements and emotions,

[0868] A means of analyzing the emotional state of the sender,

[0869] A system that includes means for providing an automated response to the caller.

[0870] (Claim 2)

[0871] The system according to claim 1, further comprising means for recording a signal and notifying the sender of that fact.

[0872] (Claim 3)

[0873] The system according to claim 1, comprising means for automatically rejecting calls based on information classified by natural language processing means and sentiment analysis means if there is a possibility of commercial activity or fraud.

[0874] "Application example 2 when combining with an emotional engine"

[0875] (Claim 1)

[0876] Means for receiving calls from communication devices,

[0877] A means of obtaining and identifying caller characteristics from a database,

[0878] A speech-to-text conversion method,

[0879] A natural language processing tool that analyzes transcribed information and classifies the requirements,

[0880] A means of analyzing the caller's emotions and deciding whether or not to transfer the call to the user based on the caller's state,

[0881] A means of providing an automated response to the caller,

[0882] A means of evaluating urgency and danger based on emotion analysis results and setting call priorities,

[0883] A system that includes this.

[0884] (Claim 2)

[0885] The system according to claim 1, further comprising means for recording a call and notifying the caller accordingly.

[0886] (Claim 3)

[0887] The system according to claim 1, comprising means for evaluating the caller's psychological state based on the emotion analysis results and, if necessary, prioritizing notification or processing of the call. [Explanation of Symbols]

[0888] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving calls from a communication terminal, A means of obtaining and identifying sender information from a database, A speech recognition method that converts speech to text, A natural language processing tool that analyzes transcribed information and classifies requirements, A means of determining whether or not to transfer a call to the user based on the caller's requirements, A system that includes means for providing an automated response to the caller.

2. The system according to claim 1, further comprising means for recording a call and notifying the caller of that fact.

3. The system according to claim 1, comprising means for automatically rejecting calls if they are potentially business or fraudulent, based on information classified by natural language processing means.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A