system

A system that converts customer conversations to text, analyzes keywords, and provides real-time product suggestions addresses the challenge of staff knowledge gaps in retail, enhancing customer service efficiency and satisfaction.

JP2026062241APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

In modern retail, new or less experienced staff struggle to efficiently and accurately provide customers with appropriate product information due to the complexity of product knowledge requirements, leading to delays and mismatches in customer service.

Method used

A system that includes a terminal to capture and encrypt customer conversations, a server to convert audio data to text, analyze keywords, and select optimal products, and a terminal to play back the information, with mode switching capabilities for staff to provide real-time product suggestions.

Benefits of technology

Enables staff to quickly and accurately recommend products that meet customer needs, improving customer satisfaction and service efficiency without relying on personal knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062241000001_ABST
    Figure 2026062241000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】means for the terminal to acquire a conversation with the user as voice data, means for the terminal to transmit the acquired voice data to the server, means for the server to convert the voice data into text data, means for the server to analyze the text data and extract keywords, means for the server to select an optimal product based on the extracted keywords, means for the server to convert the information of the selected product into voice data and transmit it to the terminal, means for the terminal to play the received voice data, means for the terminal to switch modes, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern retail, in order to improve the efficiency of customer service, it is required that staff can appropriately and quickly propose products that meet customer demands. However, in order to handle a huge amount of product information and meet the needs of different customers, high - level knowledge and experience are required for staff. Therefore, it becomes a problem that especially new staff or staff with little experience have difficulties in customer service. The purpose of the present invention is to solve this problem and improve customer service by providing a system that automatically provides optimal product information to staff from conversations with customers.

Means for Solving the Problems

[0005] The present invention includes means for a terminal to acquire conversations with a user as audio data and transmit the acquired audio data to a server. Furthermore, it includes means for the server to convert the audio data into text data and analyze the text data to extract keywords. It also includes means for the server to select the optimal product based on the extracted keywords and to convert the information of the selected product into audio data and transmit it to the terminal. The terminal includes means for playing back the received audio data and providing information to staff. The terminal also includes means for switching modes, which allows for easy switching between smartphone suggestion mode and normal intercom mode. In this way, staff can provide customers with optimal product information in real time, improving the efficiency and quality of customer service.

[0006] A "terminal" is a device that acquires conversations with a user as audio data and sends that data to a server.

[0007] A "server" is a computer system that converts audio data transmitted from a terminal into text data, analyzes it to extract keywords, and generates appropriate product information.

[0008] "Audio data" refers to data that digitally represents conversations between users (customers) and staff.

[0009] "Text data" refers to string information converted from audio data using speech recognition technology.

[0010] "Keywords" are important words or phrases extracted from text data and are analyzed to indicate customer needs and requests.

[0011] "Natural language processing technology" refers to the technology that enables computers to understand, analyze, and manipulate human language.

[0012] A "recommendation system" is an algorithm or technology that selects the most suitable product from a large number of options based on extracted keywords.

[0013] "Speech synthesis technology" is a technology that converts text information into speech data.

[0014] "Mode switching" is a function that allows a device to switch between different operating modes (for example, smartphone suggestion mode and normal mode). [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the language used in the following description will be explained.

[0018] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] The system of this invention aims to analyze conversations with customers and provide optimal product information. This system includes terminals, servers, and multiple technical means that utilize voice and text data.

[0037] Specifically, the system operates as follows:

[0038] Terminal operation

[0039] The terminal captures audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and transmitted to the server in an encrypted state.

[0040] Server Processing

[0041] The server first converts the received audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing technology to extract keywords that indicate customer needs.

[0042] Based on the extracted keywords, the server searches its internal database and selects the optimal smartphone model using a predetermined algorithm. For example, conditions such as "camera performance," "battery life," and "price range" may be extracted as keywords. In this case, smartphones with these characteristics will be selected.

[0043] Product Information

[0044] The information from the selected smartphones is converted back into audio data using speech synthesis technology on the server, encrypted, and sent to the terminal. The terminal plays this received audio data and presents the product information to the staff verbally.

[0045] Mode switching

[0046] The terminal can switch modes with the press of a button according to the staff's request. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[0047] Specific example

[0048] The following are some specific use cases.

[0049] 1. The user (customer) enters the store and begins a conversation with the staff.

[0050] User: "I'm looking for a smartphone with a great camera and long battery life."

[0051] 2. The device captures this conversation as audio data and sends it to the server.

[0052] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[0053] 4. The server searches the database for multiple smartphone models based on keywords and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[0054] 5. The server converts the selection results into audio data and sends it to the terminal.

[0055] 6. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0056] In this way, staff can quickly suggest smartphones that meet customer needs. This system makes it possible to provide high-quality customer service without relying on the staff's knowledge.

[0057] ---

[0058] The above describes specific embodiments for carrying out the present invention.

[0059] The following describes the processing flow.

[0060] Step 1:

[0061] The terminal begins voice input when a staff member starts a conversation with a customer. The conversation is captured in real time as audio data in digital format.

[0062] Step 2:

[0063] The device encrypts the acquired audio data and securely transmits it to the server.

[0064] Step 3:

[0065] The server decrypts the received encrypted audio data. Then, it uses speech recognition technology to convert the audio data into text data.

[0066] Step 4:

[0067] The server analyzes the converted text data and uses natural language processing techniques to extract important keywords. For example, keywords such as "camera performance," "battery life," and "price" may be extracted.

[0068] Step 5:

[0069] The server searches its internal database based on the extracted keywords and lists the relevant smartphone models.

[0070] Step 6:

[0071] The server uses an evaluation algorithm to select the optimal smartphone model from the listed models.

[0072] Step 7:

[0073] The server generates text data describing the features of the selected smartphone model. For example, it might include information such as, "Model X features a high-performance camera and a large-capacity battery."

[0074] Step 8:

[0075] The server converts the generated text data into speech data using speech synthesis technology. Then, it encrypts the speech data and sends it to the terminal.

[0076] Step 9:

[0077] The device decrypts the received encrypted audio data and prepares it for playback as audio.

[0078] Step 10:

[0079] The device transmits the played audio data to staff via an intercom. This allows staff to make real-time recommendations for the most suitable smartphone for each customer.

[0080] Step 11:

[0081] The terminal allows staff to easily switch between smartphone suggestion mode and normal intercom mode by operating a button on the intercom.

[0082] (Example 1)

[0083] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0084] In modern retail, responding promptly to customer needs is crucial, but not all staff possess extensive product knowledge. This can lead to delays in recommending the best product or even provide incorrect information. Furthermore, a mismatch between staff information and customer requests can reduce customer satisfaction. Additionally, there's a lack of systems to efficiently process useful information gathered during conversations and provide optimal product recommendations.

[0085] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0086] In this invention, the server includes means for converting speech into text data using a speech recognition system, means for extracting keywords from the text data using natural language processing technology, and means for searching an internal database and selecting the optimal product using an evaluation algorithm. This enables the system to accurately grasp customer requirements and make quick and accurate product suggestions without relying on the knowledge of staff.

[0087] A "terminal" is a device that acquires conversations with users as audio data and transmits it to a server.

[0088] A "server" is a computer system that receives audio data sent from a terminal, converts it to text data, analyzes it, selects products, and converts it back to audio data.

[0089] "Audio data" refers to data obtained by converting the voices of users or staff into digital signals.

[0090] "Text data" refers to character information converted from speech data using speech recognition technology.

[0091] "Keywords" are important words or phrases extracted from text data that indicate customer needs.

[0092] A "product" is a product that the server proposes based on customer needs, and in this example, it is a smartphone.

[0093] A "speech recognition system" is a system that possesses the technology to convert speech data into text data.

[0094] "Natural language processing technology" is a technique that analyzes text data and extracts keywords.

[0095] "Speech synthesis technology" is a technology that converts text data into speech data.

[0096] An "internal database" is a data store containing product information stored within the server.

[0097] An "evaluation algorithm" is a computational method for selecting the optimal product based on extracted keywords.

[0098] "Encryption" is the process of transforming data for security purposes.

[0099] "Decryption" is the process of restoring encrypted data to its original form.

[0100] The system of this invention aims to analyze conversations with customers and provide optimal product information. This system consists of a terminal, a server, and multiple technical means such as speech recognition, natural language processing, database search, and speech synthesis technology.

[0101] Terminal operation

[0102] The terminal uses a voice input device (microphone) to capture conversations with the user in real time. The captured voice data is converted into a digital signal and then encrypted. An advanced algorithm such as AES-256 is used for encryption, and the data is sent to the server via a secure communication protocol (e.g., HTTPS).

[0103] Server Processing

[0104] Speech recognition and text conversion

[0105] The server receives encrypted audio data and first decrypts it. The decrypted audio data is then input into a speech recognition system (e.g., Google® Cloud Speech-to-Text API) to convert the audio data into text data.

[0106] Text data analysis

[0107] The server analyzes the obtained text data using natural language processing techniques (e.g., spaCy, NLTK) to extract keywords that indicate customer needs. This keyword extraction clarifies the user's requirements.

[0108] Product selection

[0109] Based on the extracted keywords, the server searches its internal database. This database contains detailed product information. The search results are analyzed using a predetermined evaluation algorithm (e.g., multi-criteria evaluation method) to select the optimal product.

[0110] Product Information

[0111] Reconversion to audio data

[0112] The information on the selected products is converted back into audio data by the server using speech synthesis technology (e.g., Amazon Polly). This audio data is then encrypted again and sent to the device.

[0113] Playback of audio data

[0114] The device decrypts the received encrypted audio data and plays it back using the smartphone's speaker. Based on this audio, staff can then provide customers with the most relevant product information.

[0115] Mode switching

[0116] The terminal can be switched between smartphone suggestion mode and normal intercom mode by staff depending on the usage situation. This function improves the convenience of customer service.

[0117] Specific example

[0118] The following are some specific use cases:

[0119] 1. The user (customer) enters the store and begins a conversation with the staff.

[0120] User: "I'm looking for a smartphone with a great camera and long battery life."

[0121] 2. The device captures this conversation using its microphone, encrypts the audio data, and sends it to the server.

[0122] 3. The server receives the audio data and converts it into text data using a speech recognition system.

[0123] Example translation: "I'm looking for a smartphone with a great camera and long battery life."

[0124] 4. The server extracts the keywords "camera performance" and "long battery life" from the retrieved text.

[0125] 5. The server searches the database and determines that "Model X" is the optimal model based on its evaluation algorithm.

[0126] Example of selection reason: Model X has a high-performance camera and a large-capacity battery.

[0127] 6. The server converts the information from Model X back into audio data, encrypts it, and transmits it.

[0128] 7. The device decodes the audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0129] Examples of prompt statements

[0130] The following are examples of prompts to input into the generated AI model:

[0131] "Convert this audio data into text data, extract keywords that align with the customer's needs, and suggest the most suitable smartphone model."

[0132] Thus, the system of the present invention enables staff to quickly propose the most suitable products that meet customer needs. This leads to improved customer satisfaction and more efficient customer service.

[0133] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0134] Step 1:

[0135] The device acquires conversations with the user in real time. The voice input device (microphone) collects the audio signal and converts it into a digital format. The input is the user's spoken voice, and the output is audio data in digital format.

[0136] Step 2:

[0137] The terminal encrypts the acquired digital audio data using an advanced encryption algorithm (e.g., AES-256) and sends it to the server using a secure communication protocol (e.g., HTTPS). The input is digital audio data, and the output is encrypted audio data.

[0138] Step 3:

[0139] The server receives encrypted audio data and first decrypts it. The input is the encrypted audio data, and the output is the decrypted audio data. Specifically, the AES-256 decryption algorithm is used.

[0140] Step 4:

[0141] The server inputs the decoded audio data into a speech recognition system (e.g., Google Cloud Speech-to-Text API) and converts the audio data into text data. The input is decoded audio data, and the output is text data. Specifically, it performs API calls.

[0142] Step 5:

[0143] The server analyzes the obtained text data using natural language processing techniques (e.g., spaCy, NLTK) and extracts keywords that indicate customer needs from the text data. The input is text data, and the output is the extracted keywords. Specifically, it performs text analysis and keyword extraction.

[0144] Step 6:

[0145] The server searches its internal database based on the extracted keywords. From the search results, it selects the optimal product using a predetermined evaluation algorithm (e.g., multi-criteria evaluation method). The input is the extracted keywords, and the output is information on the optimal product. Specifically, it performs optimization using database queries and evaluation algorithms.

[0146] Step 7:

[0147] The server converts the information of the selected products back into audio data using speech synthesis technology (e.g., Amazon Polly). The input is the optimal product information, and the output is audio data. Specifically, it makes API calls to convert text to speech.

[0148] Step 8:

[0149] The server sends encrypted audio data to the terminal. Input and output are the same as above, and encryption is performed again during the transmission process. Secure protocols such as HTTPS are used for communication.

[0150] Step 9:

[0151] The device decrypts the received encrypted audio data and plays it back using the smartphone's speaker. The input is the encrypted audio data, and the output is the played audio information. Specifically, it performs decryption and audio playback.

[0152] Step 10:

[0153] The terminal switches between smartphone suggestion mode and normal intercom mode via operation at the request of the staff. Specifically, mode switching is performed by pressing a physical button or touching the screen. The input is the staff's operation, and the output is the mode change.

[0154] (Application Example 1)

[0155] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0156] Traditional brick-and-mortar stores face the challenge of staff being unable to quickly provide customers with appropriate product information. In particular, staff with limited product knowledge struggle to recommend the most suitable products for customer needs, leading to decreased customer satisfaction. Furthermore, the lack of systems capable of analyzing real-time conversational data and providing accurate product information hindered efficient service delivery.

[0157] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0158] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and extracting keywords, means for selecting the optimal product based on the extracted keywords, and means for using a generative AI model when converting product information into voice data. This makes it possible to analyze customer interactions in real time and provide optimal product information in voice.

[0159] A "terminal" is a device that acquires conversations with users as audio data and transmits it to a server.

[0160] "Audio data" refers to data obtained by converting the audio signal of a conversation with a user into a digital format.

[0161] A "server" is a computer system that receives audio data, converts it into text data, extracts keywords, and generates optimal product information based on that data.

[0162] "Text data" refers to written data converted from audio data by the server.

[0163] "Keywords" are important words or phrases extracted from text data that indicate customer needs or requests.

[0164] "Product information" refers to detailed information about the products selected by the server based on keywords.

[0165] A "generative AI model" refers to an algorithm or model that uses artificial intelligence to generate data, and in this invention, it is used specifically for generating audio data.

[0166] "Playback" refers to the operation in which a device processes the received audio data and outputs it as audio.

[0167] "Mode switching" refers to a function that allows a terminal to select different operating modes, and in this invention, it mainly involves switching between a product suggestion mode and a normal mode.

[0168] "Encryption" is the process by which audio data is digitally transformed to protect it from unauthorized access.

[0169] "Decryption" is the process of restoring encrypted data to its original state so that it can be interpreted as correct information.

[0170] "Natural language processing technology" is artificial intelligence technology that analyzes text data to understand its meaning and structure.

[0171] A "recommendation system" is an algorithm or software that selects and recommends the most suitable products or services based on user requirements.

[0172] This invention is a system for streamlining customer service in physical stores and improving customer satisfaction, providing product information using terminals, servers, and generative AI models. This system can analyze customer conversations using voice data and provide optimal product information.

[0173] System Configuration

[0174] terminal

[0175] The terminal is a device used by store staff to converse with customers. The terminal has the function to acquire voice data in real time, which is then encrypted and sent to a server. Furthermore, the terminal plays back the voice data received from the server, providing product information to the staff. The terminal also has a mode switching function, allowing it to switch between product suggestion mode and normal mode as needed.

[0176] server

[0177] The server receives audio data transmitted from the terminal and performs multiple analysis processes. First, it converts the audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing (NLP) technology to extract keywords that indicate customer needs. Based on the extracted keywords, the server searches its internal database and selects the most suitable product using its recommendation system. In this process, it uses a generative AI model to convert product information back into audio data.

[0178] Specific processing steps

[0179] The server uses standard cloud services such as "Google Cloud Speech-to-Text" for speech recognition to convert speech data into text data. Next, it uses software such as "Google Cloud Natural Language API" and "spaCy" for natural language processing to extract keywords from the text data. Then, based on a recommendation system, it searches the database for the best products that match the keywords and converts the product information into speech data using a generative AI model (e.g., "GPT-4®").

[0180] Specific examples of actions

[0181] 1. A customer enters the store and begins a conversation with a staff member.

[0182] Customer: "I'm looking for a smartphone with a great camera and long battery life."

[0183] 2. The device captures this conversation as audio data, encrypts it, and sends it to the server.

[0184] 3. The server converts the audio data into text data and uses NLP technology to extract keywords such as "camera performance" and "long battery life."

[0185] 4. The server searches the database based on keywords and selects the most suitable product. For example, "Model X" might be selected as the appropriate smartphone.

[0186] 5. The server converts the selected product information using the generation AI model into voice data and sends it to the terminal.

[0187] 6. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0188] Example of a prompt

[0189] Customer: "I'm looking for a smartphone with a great camera and long battery life."

[0190] System: "The Model X features a high-performance camera and a large-capacity battery."

[0191] In this way, store staff can respond quickly to customer needs.

[0192] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0193] Step 1:

[0194] The device acquires conversations with users as audio data. Specifically, store staff use a device (e.g., a smartphone) to interact with customers, and the device's microphone collects audio data during this interaction. The input is the conversation between the customer and the staff, and the output is the acquired audio data.

[0195] Step 2:

[0196] The terminal encrypts the acquired audio data and sends it to the server. Specifically, the terminal secures the audio data using an encryption algorithm such as AES and sends it to the server using the HTTPS protocol. The input is the acquired audio data, and the output is the encrypted audio data.

[0197] Step 3:

[0198] The server receives and decrypts encrypted audio data. Specifically, the server applies a decryption algorithm upon receipt to generate the original audio data. This ensures secure data transfer. The input is encrypted audio data, and the output is decrypted audio data.

[0199] Step 4:

[0200] The server converts the received audio data into text data using speech recognition technology. Specifically, it uses the "Google Cloud Speech-to-Text" API to analyze the audio data and transcribe it. The input is the decoded audio data, and the output is the corresponding text data.

[0201] Step 5:

[0202] The server analyzes text data using natural language processing (NLP) techniques to extract keywords. Specifically, it uses NLP software such as "Google Cloud Natural Language API" and "spaCy" to extract key keywords that indicate customer needs from the text data. The input is text data, and the output is the extracted keywords.

[0203] Step 6:

[0204] The server searches its internal database based on the extracted keywords and selects the most suitable product. Specifically, it uses a recommendation system (e.g., a filtering algorithm) to search for product information within the database and select the product that best matches the keywords. The input is the extracted keywords, and the output is the selected product information.

[0205] Step 7:

[0206] The server converts selected product information into audio data using a generative AI model. Specifically, it uses a generative AI model such as "GPT-4" to convert the product information into text, and then uses the "Google Cloud Text-to-Speech" API to convert it into audio data. The input is selected product information, and the output is audio data.

[0207] Step 8:

[0208] The server sends the generated audio data to the terminal. Specifically, it re-encrypts the audio data and sends it to the terminal via the HTTPS protocol. The input is the generated audio data, and the output is the encrypted audio data.

[0209] Step 9:

[0210] The terminal decrypts and plays back the received audio data. Specifically, it decrypts the received encrypted data and provides product information to staff via audio using the terminal's speaker. The input is the received encrypted audio data, and the output is the played audio information.

[0211] This series of processing steps enables store staff to provide customers with product information quickly and accurately in response to their needs.

[0212] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0213] The present invention aims to analyze customer conversations and provide optimal product information. This system includes terminals, servers, an emotion engine, and multiple technical means utilizing voice and text data.

[0214] Specifically, the system operates as follows:

[0215] Terminal operation

[0216] The terminal captures audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and transmitted to the server in an encrypted state.

[0217] Server Processing

[0218] The server first converts the received audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing technology to extract keywords that indicate customer needs.

[0219] How the emotion engine works

[0220] Furthermore, the voice data is analyzed by an emotion engine to recognize the customer's emotional state (e.g., satisfaction, dissatisfaction, excitement, questioning, etc.). Based on this emotional information, the emotion engine adjusts the importance and priority of keywords in product recommendations.

[0221] Selection of the optimal product

[0222] The server searches its internal database by combining extracted keywords and sentiment information, and selects the optimal smartphone model using a predetermined algorithm. For example, conditions such as "camera performance," "battery life," and "price range" may be extracted as keywords. If the sentiment engine recognizes that the customer is excited, it may make suggestions, including bold promotions.

[0223] Product Information

[0224] The information from the selected smartphones is converted back into audio data using speech synthesis technology on the server, encrypted, and sent to the terminal. The terminal plays this received audio data and presents the product information to the staff verbally.

[0225] The emotion engine adjusts its tone according to the user's emotions; if the user appears anxious, it adjusts to provide information in a calmer voice. This is expected to further enhance user satisfaction.

[0226] Mode switching

[0227] The terminal can switch modes with the press of a button according to the staff's request. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[0228] Specific example

[0229] The following are some specific use cases.

[0230] 1. The user (customer) enters the store and begins a conversation with the staff.

[0231] User: "I'm looking for a smartphone with a great camera and long battery life."

[0232] 2. The device captures this conversation as audio data and sends it to the server.

[0233] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[0234] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[0235] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[0236] 6. The server converts the selection results into audio data and sends it to the terminal. At this time, the emotion engine applies a tone that corresponds to the user's excitement level.

[0237] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0238] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[0239] In this way, staff can quickly and accurately suggest smartphones that match the customer's needs and feelings. This system makes it possible to provide high-quality customer service without relying on the staff's knowledge.

[0240] ---

[0241] The above describes specific embodiments for carrying out the present invention.

[0242] The following describes the processing flow.

[0243] Step 1:

[0244] The terminal begins voice input when a staff member starts a conversation with a customer. The conversation is captured in real time as audio data in digital format.

[0245] Step 2:

[0246] The device encrypts the acquired audio data and securely transmits it to the server.

[0247] Step 3:

[0248] The server decrypts the received encrypted audio data. Then, it uses speech recognition technology to convert the audio data into text data.

[0249] Step 4:

[0250] The server analyzes the converted text data and uses natural language processing techniques to extract important keywords. For example, keywords such as "camera performance," "battery life," and "price" may be extracted.

[0251] Step 5:

[0252] The server simultaneously passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine recognizes the customer's emotions, such as excitement, satisfaction, dissatisfaction, or questioning, based on their voice tone and speaking style.

[0253] Step 6:

[0254] The server searches its internal database based on the extracted keywords and sentiment information obtained from the sentiment engine, and lists the relevant smartphone models.

[0255] Step 7:

[0256] The server selects the optimal smartphone model from the listed models using an evaluation algorithm that combines emotional information. For example, if the customer appears anxious, a model emphasizing ease of use will be selected.

[0257] Step 8:

[0258] The server generates information about the selected smartphone model as text data. For example, it might include information such as, "Model X features a high-performance camera and a large-capacity battery."

[0259] Step 9:

[0260] The server converts the generated text data into speech data using speech synthesis technology. During this process, an emotion engine adjusts the tone of the speech. For example, if the user is excited, the speech will be generated in an energetic tone; if they are calm, it will be generated in a gentle tone.

[0261] Step 10:

[0262] The server encrypts the audio data and sends it to the terminal.

[0263] Step 11:

[0264] The device decrypts the received encrypted audio data and prepares it for playback.

[0265] Step 12:

[0266] The device transmits the played audio data to staff via an intercom. This allows staff to make real-time recommendations for the most suitable smartphone for each customer.

[0267] Step 13:

[0268] The terminal allows staff to easily switch between smartphone suggestion mode and normal intercom mode by operating a button on the intercom.

[0269] Specific example:

[0270] 1. The user (customer) enters the store and begins a conversation with the staff.

[0271] User: "I'm looking for a smartphone with a great camera and long battery life."

[0272] 2. The device captures this conversation as audio data and sends it to the server.

[0273] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[0274] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[0275] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[0276] 6. The server converts the selection results into audio data, applies a tone corresponding to the user's excitement level using the emotion engine, and sends it to the terminal.

[0277] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0278] 8. The terminal can be switched from smartphone suggestion mode to normal intercom mode by the staff operating a button on the intercom.

[0279] In this way, the staff can quickly and accurately propose a smartphone that meets the needs and feelings of customers. This system enables the provision of high-quality customer service without relying on the knowledge of the staff.

[0280] (Example 2)

[0281] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart device 14 is referred to as a "terminal".

[0282] In the conventional customer response system, product proposals and responses that rely on the knowledge of the staff are mainstream, and it is difficult to provide high-quality services that respond immediately to the needs and feelings of customers. In addition, manually performing text conversion and sentiment analysis of voice data takes time and lacks accuracy. This may lead to a decrease in customer satisfaction and have an adverse impact on the sales of the store. Furthermore, since security measures in the transmission and reception of voice data are not sufficient, there is a risk of information leakage.

[0283] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0284] In this invention, the server includes means for converting voice data into text data using voice recognition technology, means for analyzing the text data using natural language processing technology and extracting keywords, and means for analyzing the voice data by a sentiment analysis engine to recognize the emotional state of the customer. This enables quick response to the needs and feelings of customers and the provision of high-quality and secure customer service.

[0285] The "terminal" is a device that acquires conversations between the staff and customers and transmits voice data to the server.

[0286] The "server" is a central processing unit that receives voice data, performs analysis, and provides optimal product information.

[0287] "Voice data" refers to digital signals obtained by a terminal and including the conversation content between staff and customers.

[0288] "Encryption" refers to the process of converting and protecting voice data for security purposes.

[0289] "Decryption" is the process of restoring encrypted voice data to its original state.

[0290] "Voice recognition technology" is the technology for converting voice data into text data.

[0291] "Text data" is data representing conversation content as character information, which is converted by voice recognition technology.

[0292] "Natural language processing technology" is a general term for technologies that analyze text data and extract keywords.

[0293] "Keyword" is an important phrase indicating the needs of customers.

[0294] "Emotion engine" is a system that analyzes voice data and recognizes the emotional state of customers.

[0295] "Emotional state" refers to the mental states such as "satisfaction", "dissatisfaction", "excitement", "doubt", etc. shown by customers during conversations.

[0296] "Database" is a data storage system for storing product information and other necessary information.

[0297] "Algorithm" refers to the calculation procedures or processing procedures for selecting the optimal product based on the extracted keywords and emotional information.

[0298] "Voice synthesis technology" is the technology for converting text data into voice data.

[0299] "Playback" refers to the process of outputting the audio data received by the terminal as audio.

[0300] "Mode" indicates the operating state of the terminal and includes, for example, the smartphone proposal mode and the normal incoming mode.

[0301] The system of the present invention is constructed to provide optimal product information through interaction with users (customers). This system includes a terminal, a server, an emotion engine, and multiple technical means utilizing audio and text data.

[0302] Specifically, the system operates using the following hardware and software.

[0303] Operation of the Terminal

[0304] When a store staff member converses with a customer, the terminal acquires the audio in real time. The hardware used includes a high-quality microphone with a noise cancellation function. The acquired audio data is converted into digital format using digital signal processing technology and transmitted to the server in a state where security is enhanced with AES-256 encryption technology.

[0305] Server Processing

[0306] The server receives the encrypted audio data transmitted from the terminal and first decrypts it. Then, it converts the audio data into text data using the Google Speech-to-Text API. In this conversion process, noise is removed and the quality is maintained. Next, it analyzes the text data using natural language processing libraries (e.g., spaCy or NLTK) and extracts keywords indicating the customer's needs. The extracted keywords specifically indicate the product features required by the customer.

[0307] Operation of the Emotion Engine

[0308] The server has a built-in emotion engine (e.g., IBM Watson® Tone Analyzer) that analyzes the customer's emotional state based on voice data. The emotion engine analyzes the customer's voice tone, speed, pitch, etc., and determines emotional states such as "satisfied," "dissatisfied," "excited," and "questionable." This emotional information is used to adjust the importance and priority of keywords in the product proposal process described later.

[0309] Selection of the optimal product

[0310] The server searches its internal database based on extracted keywords and sentiment information, and selects the optimal product using a pre-configured algorithm (e.g., weighting calculations and evaluation functions). For example, if a customer prioritizes "camera performance" and "battery life," and the sentiment engine detects excitement, a product proposal, including a bold promotion, will be made.

[0311] Product Information

[0312] Information on selected products is converted back into audio data by the server using the Google Text-to-Speech API, and an emotion engine sets a tone that matches the customer's emotions. The re-encrypted audio data is sent to the terminal. The terminal decrypts the received audio data and plays it back through the staff member's earphones or speaker, providing the product information. Text information is also displayed on the terminal's screen.

[0313] Mode switching

[0314] The terminal features an interface that can be operated according to staff requests, and allows for easy switching between smartphone suggestion mode and normal intercom mode.

[0315] Specific example

[0316] The following are some specific use cases.

[0317] 1. The user (customer) enters the store and begins a conversation with the staff.

[0318] User: "I'm looking for a smartphone with a great camera and long battery life."

[0319] 2. The device captures this conversation as audio data and sends it to the server. This is done using the built-in microphone to capture high-quality audio data, which is then encrypted.

[0320] 3. The server converts the audio data into text data using the Google Speech-to-Text API and performs noise filtering. Next, it uses natural language processing technology (e.g., spaCy) to extract keywords such as "camera performance" and "long battery life."

[0321] 4. An emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes that the user is in an excited state.

[0322] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm (e.g., a weighting algorithm). For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[0323] 6. The server converts the selection results into audio data using the Google Text-to-Speech API and sends it to the device. At this time, the emotion engine applies a tone that corresponds to the user's level of excitement.

[0324] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0325] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[0326] Examples of prompts for generative AI models

[0327] The following are examples of prompts for a generative AI model.

[0328] User: "I'm looking for a smartphone with a great camera and long battery life."

[0329] This system allows staff to quickly and accurately recommend smartphones that match the customer's needs and feelings. It also enables high-quality customer service without relying on the staff's knowledge.

[0330] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0331] Program processing flow

[0332] Step 1: The device acquires the audio.

[0333] The device captures conversations between staff and users (customers) in real time using its built-in microphone. The audio data is converted into a digital signal and encrypted at high speed.

[0334] Input: Audio of a conversation between staff and customer

[0335] Processing: Digital signal conversion and encryption of audio (using AES-256 encryption technology)

[0336] Output: Encrypted digital audio data

[0337] Step 2: The device sends the audio data to the server.

[0338] The device sends encrypted voice data to the server. This communication is conducted using a secure protocol.

[0339] Input: Encrypted digital audio data

[0340] Processing: Data transmission (using a secure communication protocol)

[0341] Output: Data arrives on the server.

[0342] Step 3: The server decrypts the audio data.

[0343] The server receives encrypted audio data sent from the terminal and decrypts it using AES-256 decryption technology.

[0344] Input: Encrypted digital audio data

[0345] Processing: Decoding of audio data (using AES-256 decoding technology)

[0346] Output: Decoded digital audio data

[0347] Step 4: The server converts the audio to text.

[0348] The server uses the Google Speech-to-Text API to convert the audio data into text data.

[0349] Input: Decoded digital audio data

[0350] Processing: Speech recognition (using Google Speech-to-Text API)

[0351] Output: Text data

[0352] Step 5: The server extracts the keywords.

[0353] The server uses natural language processing tools such as spaCy and NLTK to extract important keywords from text data.

[0354] Input: Text data

[0355] Processing: Natural language processing (keyword extraction, TF-IDF calculation, etc.)

[0356] Output: Extracted keywords

[0357] Step 6: The emotion engine analyzes emotions.

[0358] The server uses an emotion engine to analyze voice data and recognize the user's emotional state.

[0359] Input: Decoded digital audio data

[0360] Processing: Sentiment analysis (analyzes voice tone, word choice, speed, etc.)

[0361] Output: Emotional information (satisfied, dissatisfied, excited, questioned, etc.)

[0362] Step 7: Select the optimal product for the server.

[0363] The server searches its internal database based on the extracted keywords and sentiment information to select the most suitable product.

[0364] Input: Extracted keywords, sentiment information

[0365] Processing: Database search and evaluation algorithms (weighting calculation, similarity calculation, etc.)

[0366] Output: Optimal product information

[0367] Step 8: The server generates the audio data.

[0368] The server uses the Google Text-to-Speech API to convert selected product information into audio data. The tone is then adjusted by an emotion engine.

[0369] Input: Optimal product information

[0370] Processing: Text-to-speech (using Google Text-to-Speech API)

[0371] Output: Audio data

[0372] Step 9: The server sends encrypted audio data to the device.

[0373] The server re-encrypts the generated audio data and sends it to the terminal.

[0374] Input: Audio data

[0375] Processing: Encryption of audio data (using AES-256 encryption technology) and transmission.

[0376] Output: Encrypted audio data

[0377] Step 10: The device plays the audio data.

[0378] The terminal decodes the received audio data and plays it back through earphones or speakers worn by the staff. It also displays text information on its screen.

[0379] Input: Encrypted audio data

[0380] Processing: Decoding and playback of audio data, displaying text on the screen.

[0381] Output: Played audio, displayed text information

[0382] Step 11: The device switches modes.

[0383] The terminal's mode can be switched by staff operation. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[0384] Input: Staff operation (button press)

[0385] Processing: Mode switching

[0386] Output: New mode setting

[0387] Example of a prompt

[0388] The following are examples of prompts for a generative AI model.

[0389] User: "I'm looking for a smartphone with a great camera and long battery life."

[0390] (Application Example 2)

[0391] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0392] Conventional customer service systems often rely on subjective judgments from staff when recommending products, failing to adequately reflect customers' specific needs and emotional states. Furthermore, providing timely and appropriate product information is difficult, resulting in a lack of effective means to improve customer satisfaction. This invention aims to solve these problems by providing a system that delivers optimal product information in real time, based on customer needs and emotions.

[0393] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's emotional state using an emotion engine and adjusting the tone of product suggestions based on the emotional state; means for generating the tone of product suggestions as prompt sentences using a generation AI model; and means for extracting keywords using natural language processing technology and selecting the optimal product using a recommendation system. This makes it possible to make product suggestions in an appropriate tone that takes into account the customer's emotional state, and is expected to improve customer satisfaction.

[0394] A "terminal" is a device that acquires conversations with a user as audio data, sends that data to a server, and plays back information from the server.

[0395] A "server" is a central processing unit that converts audio data into text data, analyzes that text data to extract keywords, analyzes the user's emotional state using an emotion engine, selects the optimal product, and transmits that information to the terminal.

[0396] An "emotion engine" is an algorithm or software that can analyze a user's voice data and recognize their emotional state.

[0397] "Keywords" are important words and phrases extracted from user conversations that are useful when selecting or proposing products.

[0398] "Natural language processing technology" is a technique for analyzing text data and extracting meaningful keywords from it.

[0399] A "recommendation system" is an algorithm or software that searches a database based on extracted keywords and selects the most suitable product.

[0400] A "generative AI model" is a machine learning-based artificial intelligence model used to generate the tone of product proposals.

[0401] A "prompt sentence" is an input sentence used to generate an AI model that includes the tone of a product proposal.

[0402] "Means for switching modes" refers to methods for changing the operating mode of a device depending on its intended use, such as switching between smartphone suggestion mode and normal intercom mode.

[0403] A system for specifically implementing the present invention is described below. This system acquires conversations between customers and staff in real time, analyzes the content of those conversations and the customer's emotional state, and then provides optimal product information.

[0404] Terminal operation

[0405] The terminal first captures the audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and sent to the server in an encrypted state. Standard protocols such as SSL / TLS are used for this encryption.

[0406] Server Processing

[0407] The server first converts the received audio data into text data using speech recognition technology. Google's speech recognition API is used here. Next, this text data is analyzed using natural language processing technology to extract keywords that indicate customer needs. The Python nltk library is used for natural language processing.

[0408] Furthermore, the server uses an emotion engine to analyze the voice data and recognize the customer's emotional state (e.g., satisfaction, dissatisfaction, excitement, questioning, etc.). The emotion engine uses the Hugging Face transformers library.

[0409] Selection of the optimal product

[0410] The server searches its internal database based on extracted keywords and sentiment information, and selects the optimal product using a predetermined algorithm. During this process, the database search and recommendation system work in conjunction. The recommendation system utilizes the scikit-learn library.

[0411] The selected product information is then adjusted by an emotion engine to match the customer's emotional tone. For example, if the customer is excited, the suggestions will be presented in an energetic tone. These suggestions are generated as prompts based on a generative AI model.

[0412] Product Information

[0413] The information on the selected products is converted back into audio data by the server using speech synthesis technology. The pyttsx3 library is used for this speech synthesis. Next, the audio data is encrypted again and sent to the terminal. The terminal plays the received audio data and presents the product information to the staff verbally.

[0414] Mode switching

[0415] The device has a function that allows staff to switch modes with the press of a button, according to their request. This makes it easy to switch between smartphone suggestion mode and normal intercom mode.

[0416] Specific example

[0417] The following are some specific use cases.

[0418] 1. The user (customer) enters the store and begins a conversation with the staff.

[0419] User: "I'm looking for a smartphone with a great camera and long battery life."

[0420] 2. The device captures this conversation as audio data and sends it to the server.

[0421] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[0422] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[0423] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[0424] 6. The server converts the selection results into audio data and sends it to the terminal. At this time, the emotion engine applies a tone that corresponds to the user's excitement level.

[0425] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0426] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[0427] Example of a prompt

[0428] For example, if a customer says, "I'm looking for a smartphone with a great camera and long battery life," the following prompt will be generated.

[0429] A customer says, "I'm looking for a smartphone with excellent camera performance and long battery life." Please recommend the best product for them. Use sentiment analysis to ensure your recommendation is delivered in an appropriate tone.

[0430] This system allows staff to quickly and accurately suggest products that match the customer's needs and feelings. This is expected to improve customer satisfaction.

[0431] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0432] Step 1:

[0433] The terminal acquires the conversation with the user as audio data in real time. The input is the audio data of the conversation between the user and the staff, and the output is the audio data converted into a digital signal.

[0434] Step 2:

[0435] The terminal encrypts the acquired audio data and sends it to the server. The input is audio data converted into a digital signal, and the output is encrypted audio data. Standard protocols such as SSL / TLS are used for encryption.

[0436] Step 3:

[0437] The server converts received audio data into text data using speech recognition technology. The input is encrypted audio data, and the output is text data. Google's speech recognition API is used for speech recognition.

[0438] Step 4:

[0439] The server analyzes text data using natural language processing techniques and extracts keywords. The input is text data obtained through speech recognition technology, and the output is the analyzed keywords. The NLTK library is used for natural language processing.

[0440] Step 5:

[0441] The server analyzes voice data using an emotion engine to recognize the user's emotional state. The input is voice data acquired in real time, and the output is the recognized emotional state. The Hugging Face transformers library is used as the emotion engine.

[0442] Step 6:

[0443] The server combines extracted keywords and sentiment information to search an internal database and selects the optimal product using an evaluation algorithm. The input is keywords and sentiment information, and the output is information on the selected product. The recommendation system uses the scikit-learn library.

[0444] Step 7:

[0445] The server uses a generative AI model to adjust the tone of product suggestions and generate prompt sentences. The input is selected product information and sentiment information, and the output is the prompt sentence. A pre-trained language model is used for the generative AI model.

[0446] Step 8:

[0447] The server converts product information back into audio data, encrypts it, and sends it to the terminal. The input is a prompt message, and the output is encrypted audio data. The pyttsx3 library is used for speech synthesis.

[0448] Step 9:

[0449] The terminal decrypts and plays back the received audio data. The input is encrypted audio data, and the output is the decrypted and played audio information. This audio information is presented to staff as a product suggestion via the terminal.

[0450] Step 10:

[0451] Staff members can switch between smartphone suggestion mode and normal intercom mode by operating a button on the device. The input is the button operation signal, and the output is the mode switching status of the device.

[0452] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0453] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0454] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0455] [Second Embodiment]

[0456] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0457] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0458] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0459] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0460] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0461] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0462] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0463] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0464] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0465] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0466] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0467] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0468] The system of this invention aims to analyze conversations with customers and provide optimal product information. This system includes terminals, servers, and multiple technical means that utilize voice and text data.

[0469] Specifically, the system operates as follows:

[0470] Terminal operation

[0471] The terminal captures audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and transmitted to the server in an encrypted state.

[0472] Server Processing

[0473] The server first converts the received audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing technology to extract keywords that indicate customer needs.

[0474] Based on the extracted keywords, the server searches its internal database and selects the optimal smartphone model using a predetermined algorithm. For example, conditions such as "camera performance," "battery life," and "price range" may be extracted as keywords. In this case, smartphones with these characteristics will be selected.

[0475] Product Information

[0476] The information from the selected smartphones is converted back into audio data using speech synthesis technology on the server, encrypted, and sent to the terminal. The terminal plays this received audio data and presents the product information to the staff verbally.

[0477] Mode switching

[0478] The terminal can switch modes with the press of a button according to the staff's request. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[0479] Specific example

[0480] The following are some specific use cases.

[0481] 1. The user (customer) enters the store and begins a conversation with the staff.

[0482] User: "I'm looking for a smartphone with a great camera and long battery life."

[0483] 2. The device captures this conversation as audio data and sends it to the server.

[0484] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[0485] 4. The server searches the database for multiple smartphone models based on keywords and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[0486] 5. The server converts the selection results into audio data and sends it to the terminal.

[0487] 6. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0488] In this way, staff can quickly suggest smartphones that meet customer needs. This system makes it possible to provide high-quality customer service without relying on the staff's knowledge.

[0489] ---

[0490] The above describes specific embodiments for carrying out the present invention.

[0491] The following describes the processing flow.

[0492] Step 1:

[0493] The terminal begins voice input when a staff member starts a conversation with a customer. The conversation is captured in real time as audio data in digital format.

[0494] Step 2:

[0495] The device encrypts the acquired audio data and securely transmits it to the server.

[0496] Step 3:

[0497] The server decrypts the received encrypted audio data. Then, it uses speech recognition technology to convert the audio data into text data.

[0498] Step 4:

[0499] The server analyzes the converted text data and uses natural language processing techniques to extract important keywords. For example, keywords such as "camera performance," "battery life," and "price" may be extracted.

[0500] Step 5:

[0501] The server searches its internal database based on the extracted keywords and lists the relevant smartphone models.

[0502] Step 6:

[0503] The server uses an evaluation algorithm to select the optimal smartphone model from the listed models.

[0504] Step 7:

[0505] The server generates text data describing the features of the selected smartphone model. For example, it might include information such as, "Model X features a high-performance camera and a large-capacity battery."

[0506] Step 8:

[0507] The server converts the generated text data into speech data using speech synthesis technology. Then, it encrypts the speech data and sends it to the terminal.

[0508] Step 9:

[0509] The device decrypts the received encrypted audio data and prepares it for playback as audio.

[0510] Step 10:

[0511] The device transmits the played audio data to staff via an intercom. This allows staff to make real-time recommendations for the most suitable smartphone for each customer.

[0512] Step 11:

[0513] The terminal allows staff to easily switch between smartphone suggestion mode and normal intercom mode by operating a button on the intercom.

[0514] (Example 1)

[0515] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0516] In modern retail, responding promptly to customer needs is crucial, but not all staff possess extensive product knowledge. This can lead to delays in recommending the best product or even provide incorrect information. Furthermore, a mismatch between staff information and customer requests can reduce customer satisfaction. Additionally, there's a lack of systems to efficiently process useful information gathered during conversations and provide optimal product recommendations.

[0517] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0518] In this invention, the server includes means for converting speech into text data using a speech recognition system, means for extracting keywords from the text data using natural language processing technology, and means for searching an internal database and selecting the optimal product using an evaluation algorithm. This enables the system to accurately grasp customer requirements and make quick and accurate product suggestions without relying on the knowledge of staff.

[0519] A "terminal" is a device that acquires conversations with users as audio data and transmits it to a server.

[0520] A "server" is a computer system that receives audio data sent from a terminal, converts it to text data, analyzes it, selects products, and converts it back to audio data.

[0521] "Audio data" refers to data obtained by converting the voices of users or staff into digital signals.

[0522] "Text data" refers to character information converted from speech data using speech recognition technology.

[0523] "Keywords" are important words or phrases extracted from text data that indicate customer needs.

[0524] A "product" is a product that the server proposes based on customer needs, and in this example, it is a smartphone.

[0525] A "speech recognition system" is a system that possesses the technology to convert speech data into text data.

[0526] "Natural language processing technology" is a technique that analyzes text data and extracts keywords.

[0527] "Speech synthesis technology" is a technology that converts text data into speech data.

[0528] An "internal database" is a data store containing product information stored within the server.

[0529] An "evaluation algorithm" is a computational method for selecting the optimal product based on extracted keywords.

[0530] "Encryption" is the process of transforming data for security purposes.

[0531] "Decryption" is the process of restoring encrypted data to its original form.

[0532] The system of this invention aims to analyze conversations with customers and provide optimal product information. This system consists of a terminal, a server, and multiple technical means such as speech recognition, natural language processing, database search, and speech synthesis technology.

[0533] Terminal operation

[0534] The terminal uses a voice input device (microphone) to capture conversations with the user in real time. The captured voice data is converted into a digital signal and then encrypted. An advanced algorithm such as AES-256 is used for encryption, and the data is sent to the server via a secure communication protocol (e.g., HTTPS).

[0535] Server Processing

[0536] Speech recognition and text conversion

[0537] The server receives the encrypted audio data and first decrypts it. The decrypted audio data is then input into a speech recognition system (e.g., Google Cloud Speech-to-Text API) to convert the audio data into text data.

[0538] Text data analysis

[0539] The server analyzes the obtained text data using natural language processing techniques (e.g., spaCy, NLTK) to extract keywords that indicate customer needs. This keyword extraction clarifies the user's requirements.

[0540] Product selection

[0541] Based on the extracted keywords, the server searches its internal database. This database contains detailed product information. The search results are analyzed using a predetermined evaluation algorithm (e.g., multi-criteria evaluation method) to select the optimal product.

[0542] Product Information

[0543] Reconversion to audio data

[0544] The information on the selected products is converted back into audio data by the server using speech synthesis technology (e.g., Amazon Polly). This audio data is then encrypted again and sent to the device.

[0545] Playback of audio data

[0546] The device decrypts the received encrypted audio data and plays it back using the smartphone's speaker. Based on this audio, staff can then provide customers with the most relevant product information.

[0547] Mode switching

[0548] The terminal can be switched between smartphone suggestion mode and normal intercom mode by staff depending on the usage situation. This function improves the convenience of customer service.

[0549] Specific example

[0550] The following are some specific use cases:

[0551] 1. The user (customer) enters the store and begins a conversation with the staff.

[0552] User: "I'm looking for a smartphone with a great camera and long battery life."

[0553] 2. The device captures this conversation using its microphone, encrypts the audio data, and sends it to the server.

[0554] 3. The server receives the audio data and converts it into text data using a speech recognition system.

[0555] Example translation: "I'm looking for a smartphone with a great camera and long battery life."

[0556] 4. The server extracts the keywords "camera performance" and "long battery life" from the retrieved text.

[0557] 5. The server searches the database and determines that "Model X" is the optimal model based on its evaluation algorithm.

[0558] Example of selection reason: Model X has a high-performance camera and a large-capacity battery.

[0559] 6. The server converts the information from Model X back into audio data, encrypts it, and transmits it.

[0560] 7. The device decodes the audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0561] Examples of prompt statements

[0562] The following are examples of prompts to input into the generated AI model:

[0563] "Convert this audio data into text data, extract keywords that align with the customer's needs, and suggest the most suitable smartphone model."

[0564] Thus, the system of the present invention enables staff to quickly propose the most suitable products that meet customer needs. This leads to improved customer satisfaction and more efficient customer service.

[0565] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0566] Step 1:

[0567] The device acquires conversations with the user in real time. The voice input device (microphone) collects the audio signal and converts it into a digital format. The input is the user's spoken voice, and the output is audio data in digital format.

[0568] Step 2:

[0569] The terminal encrypts the acquired digital audio data using an advanced encryption algorithm (e.g., AES-256) and sends it to the server using a secure communication protocol (e.g., HTTPS). The input is digital audio data, and the output is encrypted audio data.

[0570] Step 3:

[0571] The server receives encrypted audio data and first decrypts it. The input is the encrypted audio data, and the output is the decrypted audio data. Specifically, the AES-256 decryption algorithm is used.

[0572] Step 4:

[0573] The server inputs the decoded audio data into a speech recognition system (e.g., Google Cloud Speech-to-Text API) and converts the audio data into text data. The input is decoded audio data, and the output is text data. Specifically, it performs API calls.

[0574] Step 5:

[0575] The server analyzes the obtained text data using natural language processing techniques (e.g., spaCy, NLTK) and extracts keywords that indicate customer needs from the text data. The input is text data, and the output is the extracted keywords. Specifically, it performs text analysis and keyword extraction.

[0576] Step 6:

[0577] The server searches its internal database based on the extracted keywords. From the search results, it selects the optimal product using a predetermined evaluation algorithm (e.g., multi-criteria evaluation method). The input is the extracted keywords, and the output is information on the optimal product. Specifically, it performs optimization using database queries and evaluation algorithms.

[0578] Step 7:

[0579] The server converts the information of the selected products back into audio data using speech synthesis technology (e.g., Amazon Polly). The input is the optimal product information, and the output is audio data. Specifically, it makes API calls to convert text to speech.

[0580] Step 8:

[0581] The server sends encrypted audio data to the terminal. Input and output are the same as above, and encryption is performed again during the transmission process. Secure protocols such as HTTPS are used for communication.

[0582] Step 9:

[0583] The device decrypts the received encrypted audio data and plays it back using the smartphone's speaker. The input is the encrypted audio data, and the output is the played audio information. Specifically, it performs decryption and audio playback.

[0584] Step 10:

[0585] The terminal switches between smartphone suggestion mode and normal intercom mode via operation at the request of the staff. Specifically, mode switching is performed by pressing a physical button or touching the screen. Input is the staff's operation, and output is the mode change.

[0586] (Application Example 1)

[0587] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0588] Traditional brick-and-mortar stores face the challenge of staff being unable to quickly provide customers with appropriate product information. In particular, staff with limited product knowledge struggle to recommend the most suitable products for customer needs, leading to decreased customer satisfaction. Furthermore, the lack of systems capable of analyzing real-time conversational data and providing accurate product information hindered efficient service delivery.

[0589] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0590] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and extracting keywords, means for selecting the optimal product based on the extracted keywords, and means for using a generative AI model when converting product information into voice data. This makes it possible to analyze customer interactions in real time and provide optimal product information in voice.

[0591] A "terminal" is a device that acquires conversations with users as audio data and transmits it to a server.

[0592] "Audio data" refers to data obtained by converting the audio signal of a conversation with a user into a digital format.

[0593] A "server" is a computer system that receives audio data, converts it into text data, extracts keywords, and generates optimal product information based on that data.

[0594] "Text data" refers to written data converted from audio data by the server.

[0595] "Keywords" are important words or phrases extracted from text data that indicate customer needs or requests.

[0596] "Product information" refers to detailed information about the products selected by the server based on keywords.

[0597] A "generative AI model" refers to an algorithm or model that uses artificial intelligence to generate data, and in this invention, it is used specifically for generating audio data.

[0598] "Playback" refers to the operation in which a device processes the received audio data and outputs it as audio.

[0599] "Mode switching" refers to a function that allows a terminal to select different operating modes, and in this invention, it mainly involves switching between a product suggestion mode and a normal mode.

[0600] "Encryption" is the process by which audio data is digitally transformed to protect it from unauthorized access.

[0601] "Decryption" is the process of restoring encrypted data to its original state so that it can be interpreted as correct information.

[0602] "Natural language processing technology" is artificial intelligence technology that analyzes text data to understand its meaning and structure.

[0603] A "recommendation system" is an algorithm or software that selects and recommends the most suitable products or services based on user requirements.

[0604] This invention is a system for streamlining customer service in physical stores and improving customer satisfaction, providing product information using terminals, servers, and generative AI models. This system can analyze customer conversations using voice data and provide optimal product information.

[0605] System Configuration

[0606] terminal

[0607] The terminal is a device used by store staff to converse with customers. The terminal has the function to acquire voice data in real time, which is then encrypted and sent to a server. Furthermore, the terminal plays back the voice data received from the server, providing product information to the staff. The terminal also has a mode switching function, allowing it to switch between product suggestion mode and normal mode as needed.

[0608] server

[0609] The server receives audio data transmitted from the terminal and performs multiple analysis processes. First, it converts the audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing (NLP) technology to extract keywords that indicate customer needs. Based on the extracted keywords, the server searches its internal database and selects the most suitable product using its recommendation system. In this process, it uses a generative AI model to convert product information back into audio data.

[0610] Specific processing steps

[0611] The server uses standard cloud services such as "Google Cloud Speech-to-Text" for speech recognition to convert speech data into text data. Next, it uses software such as "Google Cloud Natural Language API" or "spaCy" for natural language processing to extract keywords from the text data. Then, based on a recommendation system, it searches the database for the best products that match the keywords and uses a generative AI model (e.g., "GPT-4") to convert the product information into speech data.

[0612] Specific examples of actions

[0613] 1. A customer enters the store and begins a conversation with a staff member.

[0614] Customer: "I'm looking for a smartphone with a great camera and long battery life."

[0615] 2. The device captures this conversation as audio data, encrypts it, and sends it to the server.

[0616] 3. The server converts the audio data into text data and uses NLP technology to extract keywords such as "camera performance" and "long battery life."

[0617] 4. The server searches the database based on keywords and selects the most suitable product. For example, "Model X" might be selected as the appropriate smartphone.

[0618] 5. The server converts the selected product information using the generation AI model into voice data and sends it to the terminal.

[0619] 6. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0620] Example of a prompt

[0621] Customer: "I'm looking for a smartphone with a great camera and long battery life."

[0622] System: "The Model X features a high-performance camera and a large-capacity battery."

[0623] In this way, store staff can respond quickly to customer needs.

[0624] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0625] Step 1:

[0626] The device acquires conversations with users as audio data. Specifically, store staff use a device (e.g., a smartphone) to interact with customers, and the device's microphone collects audio data during this interaction. The input is the conversation between the customer and the staff, and the output is the acquired audio data.

[0627] Step 2:

[0628] The terminal encrypts the acquired audio data and sends it to the server. Specifically, the terminal secures the audio data using an encryption algorithm such as AES and sends it to the server using the HTTPS protocol. The input is the acquired audio data, and the output is the encrypted audio data.

[0629] Step 3:

[0630] The server receives and decrypts encrypted audio data. Specifically, the server applies a decryption algorithm upon receipt to generate the original audio data. This ensures secure data transfer. The input is encrypted audio data, and the output is decrypted audio data.

[0631] Step 4:

[0632] The server converts the received audio data into text data using speech recognition technology. Specifically, it uses the "Google Cloud Speech-to-Text" API to analyze the audio data and transcribe it. The input is the decoded audio data, and the output is the corresponding text data.

[0633] Step 5:

[0634] The server analyzes text data using natural language processing (NLP) techniques to extract keywords. Specifically, it uses NLP software such as "Google Cloud Natural Language API" and "spaCy" to extract key keywords that indicate customer needs from the text data. The input is text data, and the output is the extracted keywords.

[0635] Step 6:

[0636] The server searches its internal database based on the extracted keywords and selects the most suitable product. Specifically, it uses a recommendation system (e.g., a filtering algorithm) to search for product information within the database and select the product that best matches the keywords. The input is the extracted keywords, and the output is the selected product information.

[0637] Step 7:

[0638] The server converts selected product information into audio data using a generative AI model. Specifically, it uses a generative AI model such as "GPT-4" to convert the product information into text, and then uses the "Google Cloud Text-to-Speech" API to convert it into audio data. The input is selected product information, and the output is audio data.

[0639] Step 8:

[0640] The server sends the generated audio data to the terminal. Specifically, it re-encrypts the audio data and sends it to the terminal via the HTTPS protocol. The input is the generated audio data, and the output is the encrypted audio data.

[0641] Step 9:

[0642] The terminal decrypts and plays back the received audio data. Specifically, it decrypts the received encrypted data and provides product information to staff via audio using the terminal's speaker. The input is the received encrypted audio data, and the output is the played audio information.

[0643] This series of processing steps enables store staff to provide customers with product information quickly and accurately in response to their needs.

[0644] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0645] The system of this invention aims to analyze conversations with customers and provide optimal product information. This system includes terminals, servers, an emotion engine, and multiple technical means that utilize voice and text data.

[0646] Specifically, the system operates as follows:

[0647] Terminal operation

[0648] The terminal captures audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and transmitted to the server in an encrypted state.

[0649] Server Processing

[0650] The server first converts the received audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing technology to extract keywords that indicate customer needs.

[0651] How the emotion engine works

[0652] Furthermore, the voice data is analyzed by an emotion engine to recognize the customer's emotional state (e.g., satisfaction, dissatisfaction, excitement, questioning, etc.). Based on this emotional information, the emotion engine adjusts the importance and priority of keywords in product recommendations.

[0653] Selection of the optimal product

[0654] The server searches its internal database by combining extracted keywords and sentiment information, and selects the optimal smartphone model using a predetermined algorithm. For example, conditions such as "camera performance," "battery life," and "price range" may be extracted as keywords. If the sentiment engine recognizes that the customer is excited, it may make suggestions, including bold promotions.

[0655] Product Information

[0656] The information from the selected smartphones is converted back into audio data using speech synthesis technology on the server, encrypted, and sent to the terminal. The terminal plays this received audio data and presents the product information to the staff verbally.

[0657] The emotion engine adjusts its tone according to the user's emotions; if the user appears anxious, it adjusts to provide information in a calmer voice. This is expected to further enhance user satisfaction.

[0658] Mode switching

[0659] The terminal can switch modes with the press of a button according to the staff's request. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[0660] Specific example

[0661] The following are some specific use cases.

[0662] 1. The user (customer) enters the store and begins a conversation with the staff.

[0663] User: "I'm looking for a smartphone with a great camera and long battery life."

[0664] 2. The device captures this conversation as audio data and sends it to the server.

[0665] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[0666] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[0667] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[0668] 6. The server converts the selection results into audio data and sends it to the terminal. At this time, the emotion engine applies a tone that corresponds to the user's excitement level.

[0669] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0670] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[0671] In this way, staff can quickly and accurately suggest smartphones that match the customer's needs and feelings. This system makes it possible to provide high-quality customer service without relying on the staff's knowledge.

[0672] ---

[0673] The above describes specific embodiments for carrying out the present invention.

[0674] The following describes the processing flow.

[0675] Step 1:

[0676] The terminal begins voice input when a staff member starts a conversation with a customer. The conversation is captured in real time as audio data in digital format.

[0677] Step 2:

[0678] The device encrypts the acquired audio data and securely transmits it to the server.

[0679] Step 3:

[0680] The server decrypts the received encrypted audio data. Then, it uses speech recognition technology to convert the audio data into text data.

[0681] Step 4:

[0682] The server analyzes the converted text data and uses natural language processing techniques to extract important keywords. For example, keywords such as "camera performance," "battery life," and "price" may be extracted.

[0683] Step 5:

[0684] The server simultaneously passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine recognizes the customer's emotions, such as excitement, satisfaction, dissatisfaction, or questioning, based on their voice tone and speaking style.

[0685] Step 6:

[0686] The server searches its internal database based on the extracted keywords and sentiment information obtained from the sentiment engine, and lists the relevant smartphone models.

[0687] Step 7:

[0688] The server selects the optimal smartphone model from the listed models using an evaluation algorithm that combines emotional information. For example, if the customer appears anxious, a model emphasizing ease of use will be selected.

[0689] Step 8:

[0690] The server generates information about the selected smartphone model as text data. For example, it might include information such as, "Model X features a high-performance camera and a large-capacity battery."

[0691] Step 9:

[0692] The server converts the generated text data into speech data using speech synthesis technology. During this process, an emotion engine adjusts the tone of the speech. For example, if the user is excited, the speech will be generated in an energetic tone; if they are calm, it will be generated in a gentle tone.

[0693] Step 10:

[0694] The server encrypts the audio data and sends it to the terminal.

[0695] Step 11:

[0696] The device decrypts the received encrypted audio data and prepares it for playback.

[0697] Step 12:

[0698] The device transmits the played audio data to staff via an intercom. This allows staff to make real-time recommendations for the most suitable smartphone for each customer.

[0699] Step 13:

[0700] The terminal allows staff to easily switch between smartphone suggestion mode and normal intercom mode by operating a button on the intercom.

[0701] Specific example:

[0702] 1. The user (customer) enters the store and begins a conversation with the staff.

[0703] User: "I'm looking for a smartphone with a great camera and long battery life."

[0704] 2. The device captures this conversation as audio data and sends it to the server.

[0705] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[0706] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[0707] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[0708] 6. The server converts the selection results into audio data, applies a tone corresponding to the user's excitement level using the emotion engine, and sends it to the terminal.

[0709] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0710] 8. The terminal can be switched from smartphone suggestion mode to normal intercom mode by the staff operating a button on the intercom.

[0711] In this way, staff can quickly and accurately suggest smartphones that match the customer's needs and feelings. This system makes it possible to provide high-quality customer service without relying on the staff's knowledge.

[0712] (Example 2)

[0713] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0714] Traditional customer service systems rely heavily on staff knowledge for product recommendations and responses, making it difficult to provide high-quality service that responds promptly to customer needs and emotions. Furthermore, manually converting voice data to text and performing sentiment analysis is time-consuming and inaccurate. This can lead to decreased customer satisfaction and negatively impact store sales. Additionally, insufficient security measures for transmitting and receiving voice data create a risk of information leakage.

[0715] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0716] In this invention, the server includes means for converting voice data into text data using voice recognition technology, means for analyzing the text data and extracting keywords using natural language processing technology, and means for analyzing the voice data using an emotion analysis engine to recognize the customer's emotional state. This makes it possible to quickly respond to customer needs and emotions and provide high-quality, secure customer service.

[0717] A "terminal" is a device that captures conversations between staff and customers and transmits the audio data to a server.

[0718] A "server" is a central processing unit that receives and analyzes audio data and provides optimal product information.

[0719] "Voice data" refers to digital signals acquired by the terminal, including the content of conversations between staff and customers.

[0720] "Encryption" refers to the process of transforming and protecting audio data for security purposes.

[0721] "Decryption" is the process of restoring encrypted audio data to its original state.

[0722] "Speech recognition technology" is a technology that converts speech data into text data.

[0723] "Text data" refers to data that represents conversation content as written information, converted using speech recognition technology.

[0724] "Natural language processing technology" is a general term for technologies that analyze text data and extract keywords.

[0725] "Keywords" are important terms that indicate customer needs.

[0726] An "emotion engine" is a system that analyzes voice data to recognize the customer's emotional state.

[0727] "Emotional state" refers to the mental state a customer expresses during a conversation, such as "satisfaction," "dissatisfaction," "excitement," or "question."

[0728] A "database" is a data storage system used to store product information and other necessary information.

[0729] An "algorithm" refers to a set of calculation and processing steps used to select the optimal product based on extracted keywords and emotional information.

[0730] "Speech synthesis technology" is a technology that converts text data into speech data.

[0731] "Playback" is the process by which a device outputs received audio data as audio.

[0732] "Mode" refers to the operating state of the device, and includes, for example, smartphone suggestion mode and normal intercom mode.

[0733] The system of the present invention is built to provide optimal product information through interaction with users (customers). This system includes terminals, servers, an emotion engine, and multiple technical means that utilize voice and text data.

[0734] Specifically, the system operates using the following hardware and software.

[0735] Terminal operation

[0736] The terminal captures audio in real time as store staff converse with customers. The hardware used includes a high-quality microphone with noise-canceling capabilities. The captured audio data is converted to a digital format using digital signal processing technology and transmitted to the server with enhanced security using AES-256 encryption technology.

[0737] Server Processing

[0738] The server receives encrypted audio data sent from the terminal and first decrypts it. Then, it converts the audio data into text data using the Google Speech-to-Text API. This conversion process removes noise and maintains quality. Next, it analyzes the text data using a natural language processing library (e.g., spaCy or NLTK) to extract keywords that indicate customer needs. The extracted keywords specifically describe the product features that the customer is looking for.

[0739] How the emotion engine works

[0740] The server has a built-in emotion engine (e.g., IBM Watson Tone Analyzer) that analyzes the customer's emotional state based on voice data. The emotion engine analyzes the customer's voice tone, speed, pitch, etc., and determines emotional states such as "satisfied," "dissatisfied," "excited," and "questionable." This emotional information is used to adjust the importance and priority of keywords in the product proposal process described later.

[0741] Selection of the optimal product

[0742] The server searches its internal database based on extracted keywords and sentiment information, and selects the optimal product using a pre-configured algorithm (e.g., weighting calculations and evaluation functions). For example, if a customer prioritizes "camera performance" and "battery life," and the sentiment engine detects excitement, a product proposal, including a bold promotion, will be made.

[0743] Product Information

[0744] Information on selected products is converted back into audio data by the server using the Google Text-to-Speech API, and an emotion engine sets a tone that matches the customer's emotions. The re-encrypted audio data is sent to the terminal. The terminal decrypts the received audio data and plays it back through the staff member's earphones or speaker, providing the product information. Text information is also displayed on the terminal's screen.

[0745] Mode switching

[0746] The terminal features an interface that can be operated according to the staff's requests, and allows for easy switching between smartphone suggestion mode and normal intercom mode.

[0747] Specific example

[0748] The following are some specific use cases.

[0749] 1. The user (customer) enters the store and begins a conversation with the staff.

[0750] User: "I'm looking for a smartphone with a great camera and long battery life."

[0751] 2. The device captures this conversation as audio data and sends it to the server. This is done using the built-in microphone to capture high-quality audio data, which is then encrypted.

[0752] 3. The server converts the audio data into text data using the Google Speech-to-Text API and performs noise filtering. Next, it uses natural language processing technology (e.g., spaCy) to extract keywords such as "camera performance" and "long battery life."

[0753] 4. An emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes that the user is in an excited state.

[0754] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm (e.g., a weighting algorithm). For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[0755] 6. The server converts the selection results into audio data using the Google Text-to-Speech API and sends it to the device. At this time, the emotion engine applies a tone that corresponds to the user's level of excitement.

[0756] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0757] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[0758] Examples of prompts for generative AI models

[0759] The following are examples of prompts for a generative AI model.

[0760] User: "I'm looking for a smartphone with a great camera and long battery life."

[0761] This system allows staff to quickly and accurately recommend smartphones that match the customer's needs and feelings. It also enables high-quality customer service without relying on the staff's knowledge.

[0762] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0763] Program processing flow

[0764] Step 1: The device acquires the audio.

[0765] The device captures conversations between staff and users (customers) in real time using its built-in microphone. The audio data is converted into a digital signal and encrypted at high speed.

[0766] Input: Audio of a conversation between staff and customer

[0767] Processing: Digital signal conversion and encryption of audio (using AES-256 encryption technology)

[0768] Output: Encrypted digital audio data

[0769] Step 2: The device sends the audio data to the server.

[0770] The device sends encrypted voice data to the server. This communication is conducted using a secure protocol.

[0771] Input: Encrypted digital audio data

[0772] Processing: Data transmission (using a secure communication protocol)

[0773] Output: Data arrives on the server.

[0774] Step 3: The server decrypts the audio data.

[0775] The server receives encrypted audio data sent from the terminal and decrypts it using AES-256 decryption technology.

[0776] Input: Encrypted digital audio data

[0777] Processing: Decoding of audio data (using AES-256 decoding technology)

[0778] Output: Decoded digital audio data

[0779] Step 4: The server converts the audio to text.

[0780] The server uses the Google Speech-to-Text API to convert the audio data into text data.

[0781] Input: Decoded digital audio data

[0782] Processing: Speech recognition (using Google Speech-to-Text API)

[0783] Output: Text data

[0784] Step 5: The server extracts the keywords.

[0785] The server uses natural language processing tools such as spaCy and NLTK to extract important keywords from text data.

[0786] Input: Text data

[0787] Processing: Natural language processing (keyword extraction, TF-IDF calculation, etc.)

[0788] Output: Extracted keywords

[0789] Step 6: The emotion engine analyzes emotions.

[0790] The server uses an emotion engine to analyze voice data and recognize the user's emotional state.

[0791] Input: Decoded digital audio data

[0792] Processing: Sentiment analysis (analyzes voice tone, word choice, speed, etc.)

[0793] Output: Emotional information (satisfied, dissatisfied, excited, questioned, etc.)

[0794] Step 7: Select the optimal product for the server.

[0795] The server searches its internal database based on the extracted keywords and sentiment information to select the most suitable product.

[0796] Input: Extracted keywords, sentiment information

[0797] Processing: Database search and evaluation algorithms (weighting calculation, similarity calculation, etc.)

[0798] Output: Optimal product information

[0799] Step 8: The server generates the audio data.

[0800] The server uses the Google Text-to-Speech API to convert selected product information into audio data. The tone is then adjusted by an emotion engine.

[0801] Input: Optimal product information

[0802] Processing: Text-to-speech (using Google Text-to-Speech API)

[0803] Output: Audio data

[0804] Step 9: The server sends encrypted audio data to the device.

[0805] The server re-encrypts the generated audio data and sends it to the terminal.

[0806] Input: Audio data

[0807] Processing: Encryption of audio data (using AES-256 encryption technology) and transmission.

[0808] Output: Encrypted audio data

[0809] Step 10: The device plays the audio data.

[0810] The terminal decodes the received audio data and plays it back through earphones or speakers worn by the staff. It also displays text information on its screen.

[0811] Input: Encrypted audio data

[0812] Processing: Decoding and playback of audio data, displaying text on the screen.

[0813] Output: Played audio, displayed text information

[0814] Step 11: The device switches modes.

[0815] The terminal's mode can be switched by staff operation. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[0816] Input: Staff operation (button press)

[0817] Processing: Mode switching

[0818] Output: New mode setting

[0819] Example of a prompt

[0820] The following are examples of prompts for a generative AI model.

[0821] User: "I'm looking for a smartphone with a great camera and long battery life."

[0822] (Application Example 2)

[0823] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0824] Conventional customer service systems often rely on subjective judgments from staff when recommending products, failing to adequately reflect customers' specific needs and emotional states. Furthermore, providing timely and appropriate product information is difficult, resulting in a lack of effective means to improve customer satisfaction. This invention aims to solve these problems by providing a system that delivers optimal product information in real time, based on customer needs and emotions.

[0825] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's emotional state using an emotion engine and adjusting the tone of product suggestions based on the emotional state; means for generating the tone of product suggestions as prompt sentences using a generation AI model; and means for extracting keywords using natural language processing technology and selecting the optimal product using a recommendation system. This makes it possible to make product suggestions in an appropriate tone that takes into account the customer's emotional state, and is expected to improve customer satisfaction.

[0826] A "terminal" is a device that acquires conversations with a user as audio data, sends that data to a server, and plays back information from the server.

[0827] A "server" is a central processing unit that converts audio data into text data, analyzes that text data to extract keywords, analyzes the user's emotional state using an emotion engine, selects the optimal product, and transmits that information to the terminal.

[0828] An "emotion engine" is an algorithm or software that can analyze a user's voice data and recognize their emotional state.

[0829] "Keywords" are important words and phrases extracted from user conversations that are useful when selecting or proposing products.

[0830] "Natural language processing technology" is a technique for analyzing text data and extracting meaningful keywords from it.

[0831] A "recommendation system" is an algorithm or software that searches a database based on extracted keywords and selects the most suitable product.

[0832] A "generative AI model" is a machine learning-based artificial intelligence model used to generate the tone of product proposals.

[0833] A "prompt sentence" is an input sentence used to generate an AI model that includes the tone of a product proposal.

[0834] "Means for switching modes" refers to methods for changing the operating mode of a device depending on its intended use, such as switching between smartphone suggestion mode and normal intercom mode.

[0835] A system for specifically implementing the present invention is described below. This system acquires conversations between customers and staff in real time and provides optimal product information by analyzing the content of the conversation and the customer's emotional state.

[0836] Terminal operation

[0837] The terminal first captures the audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and sent to the server in an encrypted state. Standard protocols such as SSL / TLS are used for this encryption.

[0838] Server Processing

[0839] The server first converts the received audio data into text data using speech recognition technology. Google's speech recognition API is used here. Next, this text data is analyzed using natural language processing technology to extract keywords that indicate customer needs. The Python nltk library is used for natural language processing.

[0840] Furthermore, the server uses an emotion engine to analyze the voice data and recognize the customer's emotional state (e.g., satisfaction, dissatisfaction, excitement, questioning, etc.). The emotion engine uses the Hugging Face transformers library.

[0841] Selection of the optimal product

[0842] The server searches its internal database based on extracted keywords and sentiment information, and selects the optimal product using a predetermined algorithm. During this process, the database search and recommendation system work in conjunction. The recommendation system utilizes the scikit-learn library.

[0843] The selected product information is then adjusted by an emotion engine to match the customer's emotional tone. For example, if the customer is excited, the suggestions will be presented in an energetic tone. These suggestions are generated as prompts based on a generative AI model.

[0844] Product Information

[0845] The information on the selected products is converted back into audio data by the server using speech synthesis technology. The pyttsx3 library is used for this speech synthesis. Next, the audio data is encrypted again and sent to the terminal. The terminal plays the received audio data and presents the product information to the staff verbally.

[0846] Mode switching

[0847] The device has a function that allows staff to switch modes with the press of a button, according to their request. This makes it easy to switch between smartphone suggestion mode and normal intercom mode.

[0848] Specific example

[0849] The following are some specific use cases.

[0850] 1. The user (customer) enters the store and begins a conversation with the staff.

[0851] User: "I'm looking for a smartphone with a great camera and long battery life."

[0852] 2. The device captures this conversation as audio data and sends it to the server.

[0853] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[0854] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[0855] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[0856] 6. The server converts the selection results into audio data and sends it to the terminal. At this time, the emotion engine applies a tone that corresponds to the user's excitement level.

[0857] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0858] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[0859] Example of a prompt

[0860] For example, if a customer says, "I'm looking for a smartphone with a great camera and long battery life," the following prompt will be generated.

[0861] A customer says, "I'm looking for a smartphone with excellent camera performance and long battery life." Please recommend the best product for them. Use sentiment analysis to ensure your recommendation is delivered in an appropriate tone.

[0862] This system allows staff to quickly and accurately suggest products that match the customer's needs and feelings. This is expected to improve customer satisfaction.

[0863] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0864] Step 1:

[0865] The terminal acquires the conversation with the user as audio data in real time. The input is the audio data of the conversation between the user and the staff, and the output is the audio data converted into a digital signal.

[0866] Step 2:

[0867] The terminal encrypts the acquired audio data and sends it to the server. The input is audio data converted into a digital signal, and the output is encrypted audio data. Standard protocols such as SSL / TLS are used for encryption.

[0868] Step 3:

[0869] The server converts received audio data into text data using speech recognition technology. The input is encrypted audio data, and the output is text data. Google's speech recognition API is used for speech recognition.

[0870] Step 4:

[0871] The server analyzes text data using natural language processing techniques and extracts keywords. The input is text data obtained through speech recognition technology, and the output is the analyzed keywords. The NLTK library is used for natural language processing.

[0872] Step 5:

[0873] The server analyzes voice data using an emotion engine to recognize the user's emotional state. The input is voice data acquired in real time, and the output is the recognized emotional state. The Hugging Face transformers library is used as the emotion engine.

[0874] Step 6:

[0875] The server combines extracted keywords and sentiment information to search an internal database and selects the optimal product using an evaluation algorithm. The input is keywords and sentiment information, and the output is information on the selected product. The recommendation system uses the scikit-learn library.

[0876] Step 7:

[0877] The server uses a generative AI model to adjust the tone of product suggestions and generate prompt sentences. The input is selected product information and sentiment information, and the output is the prompt sentence. A pre-trained language model is used for the generative AI model.

[0878] Step 8:

[0879] The server converts product information back into audio data, encrypts it, and sends it to the terminal. The input is a prompt message, and the output is encrypted audio data. The pyttsx3 library is used for speech synthesis.

[0880] Step 9:

[0881] The terminal decrypts and plays back the received audio data. The input is encrypted audio data, and the output is the decrypted and played audio information. This audio information is presented to staff as a product suggestion via the terminal.

[0882] Step 10:

[0883] Staff members can switch between smartphone suggestion mode and normal intercom mode by operating a button on the device. The input is the button operation signal, and the output is the mode switching status of the device.

[0884] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0885] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0886] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0887] [Third Embodiment]

[0888] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0889] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0890] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0891] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0892] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0893] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0894] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0895] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0896] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0897] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0898] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0899] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0900] The system of this invention aims to analyze conversations with customers and provide optimal product information. This system includes terminals, servers, and multiple technical means that utilize voice and text data.

[0901] Specifically, the system operates as follows:

[0902] Terminal operation

[0903] The terminal captures audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and transmitted to the server in an encrypted state.

[0904] Server Processing

[0905] The server first converts the received audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing technology to extract keywords that indicate customer needs.

[0906] Based on the extracted keywords, the server searches its internal database and selects the optimal smartphone model using a predetermined algorithm. For example, conditions such as "camera performance," "battery life," and "price range" may be extracted as keywords. In this case, smartphones with these characteristics will be selected.

[0907] Product Information

[0908] The information from the selected smartphones is converted back into audio data using speech synthesis technology on the server, encrypted, and sent to the terminal. The terminal plays this received audio data and presents the product information to the staff verbally.

[0909] Mode switching

[0910] The terminal can switch modes with the press of a button according to the staff's request. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[0911] Specific example

[0912] The following are some specific use cases.

[0913] 1. The user (customer) enters the store and begins a conversation with the staff.

[0914] User: "I'm looking for a smartphone with a great camera and long battery life."

[0915] 2. The device captures this conversation as audio data and sends it to the server.

[0916] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[0917] 4. The server searches the database for multiple smartphone models based on keywords and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[0918] 5. The server converts the selection results into audio data and sends it to the terminal.

[0919] 6. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0920] In this way, staff can quickly suggest smartphones that meet customer needs. This system makes it possible to provide high-quality customer service without relying on the staff's knowledge.

[0921] ---

[0922] The above describes specific embodiments for carrying out the present invention.

[0923] The following describes the processing flow.

[0924] Step 1:

[0925] The terminal begins voice input when a staff member starts a conversation with a customer. The conversation is captured in real time as audio data in digital format.

[0926] Step 2:

[0927] The device encrypts the acquired audio data and securely transmits it to the server.

[0928] Step 3:

[0929] The server decrypts the received encrypted audio data. Then, it uses speech recognition technology to convert the audio data into text data.

[0930] Step 4:

[0931] The server analyzes the converted text data and uses natural language processing techniques to extract important keywords. For example, keywords such as "camera performance," "battery life," and "price" may be extracted.

[0932] Step 5:

[0933] The server searches its internal database based on the extracted keywords and lists the relevant smartphone models.

[0934] Step 6:

[0935] The server uses an evaluation algorithm to select the optimal smartphone model from the listed models.

[0936] Step 7:

[0937] The server generates text data describing the features of the selected smartphone model. For example, it might include information such as, "Model X features a high-performance camera and a large-capacity battery."

[0938] Step 8:

[0939] The server converts the generated text data into speech data using speech synthesis technology. Then, it encrypts the speech data and sends it to the terminal.

[0940] Step 9:

[0941] The device decrypts the received encrypted audio data and prepares it for playback as audio.

[0942] Step 10:

[0943] The device transmits the played audio data to staff via an intercom. This allows staff to make real-time recommendations for the most suitable smartphone for each customer.

[0944] Step 11:

[0945] The terminal allows staff to easily switch between smartphone suggestion mode and normal intercom mode by operating a button on the intercom.

[0946] (Example 1)

[0947] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0948] In modern retail, responding promptly to customer needs is crucial, but not all staff possess extensive product knowledge. This can lead to delays in recommending the best product or even provide incorrect information. Furthermore, a mismatch between staff information and customer requests can reduce customer satisfaction. Additionally, there's a lack of systems to efficiently process useful information gathered during conversations and provide optimal product recommendations.

[0949] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0950] In this invention, the server includes means for converting speech into text data using a speech recognition system, means for extracting keywords from the text data using natural language processing technology, and means for searching an internal database and selecting the optimal product using an evaluation algorithm. This enables the system to accurately grasp customer requirements and make quick and accurate product suggestions without relying on the knowledge of staff.

[0951] A "terminal" is a device that acquires conversations with users as audio data and transmits it to a server.

[0952] A "server" is a computer system that receives audio data sent from a terminal, converts it to text data, analyzes it, selects products, and converts it back to audio data.

[0953] "Audio data" refers to data obtained by converting the voices of users or staff into digital signals.

[0954] "Text data" refers to character information converted from speech data using speech recognition technology.

[0955] "Keywords" are important words or phrases extracted from text data that indicate customer needs.

[0956] A "product" is a product that the server proposes based on customer needs, and in this example, it is a smartphone.

[0957] A "speech recognition system" is a system that possesses the technology to convert speech data into text data.

[0958] "Natural language processing technology" is a technique that analyzes text data and extracts keywords.

[0959] "Speech synthesis technology" is a technology that converts text data into speech data.

[0960] An "internal database" is a data store containing product information stored within the server.

[0961] An "evaluation algorithm" is a computational method for selecting the optimal product based on extracted keywords.

[0962] "Encryption" is the process of transforming data for security purposes.

[0963] "Decryption" is the process of restoring encrypted data to its original form.

[0964] The system of this invention aims to analyze conversations with customers and provide optimal product information. This system consists of a terminal, a server, and multiple technical means such as speech recognition, natural language processing, database search, and speech synthesis technology.

[0965] Terminal operation

[0966] The terminal uses a voice input device (microphone) to capture conversations with the user in real time. The captured voice data is converted into a digital signal and then encrypted. An advanced algorithm such as AES-256 is used for encryption, and the data is sent to the server via a secure communication protocol (e.g., HTTPS).

[0967] Server Processing

[0968] Speech recognition and text conversion

[0969] The server receives the encrypted audio data and first decrypts it. The decrypted audio data is then input into a speech recognition system (e.g., Google Cloud Speech-to-Text API) to convert the audio data into text data.

[0970] Text data analysis

[0971] The server analyzes the obtained text data using natural language processing techniques (e.g., spaCy, NLTK) to extract keywords that indicate customer needs. This keyword extraction clarifies the user's requirements.

[0972] Product selection

[0973] Based on the extracted keywords, the server searches its internal database. This database contains detailed product information. The search results are analyzed using a predetermined evaluation algorithm (e.g., multi-criteria evaluation method) to select the optimal product.

[0974] Product Information

[0975] Reconversion to audio data

[0976] The information on the selected products is converted back into audio data by the server using speech synthesis technology (e.g., Amazon Polly). This audio data is then encrypted again and sent to the device.

[0977] Playback of audio data

[0978] The device decrypts the received encrypted audio data and plays it back using the smartphone's speaker. Based on this audio, staff can then provide customers with the most relevant product information.

[0979] Mode switching

[0980] The terminal can be switched between smartphone suggestion mode and normal intercom mode by staff depending on the usage situation. This function improves the convenience of customer service.

[0981] Specific example

[0982] The following are some specific use cases:

[0983] 1. The user (customer) enters the store and begins a conversation with the staff.

[0984] User: "I'm looking for a smartphone with a great camera and long battery life."

[0985] 2. The device captures this conversation using its microphone, encrypts the audio data, and sends it to the server.

[0986] 3. The server receives the audio data and converts it into text data using a speech recognition system.

[0987] Example translation: "I'm looking for a smartphone with a great camera and long battery life."

[0988] 4. The server extracts the keywords "camera performance" and "long battery life" from the retrieved text.

[0989] 5. The server searches the database and determines that "Model X" is the optimal model based on its evaluation algorithm.

[0990] Example of selection reason: Model X has a high-performance camera and a large-capacity battery.

[0991] 6. The server converts the information from Model X back into audio data, encrypts it, and transmits it.

[0992] 7. The device decodes the audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[0993] Examples of prompt statements

[0994] The following are examples of prompts to input into the generated AI model:

[0995] "Convert this audio data into text data, extract keywords that align with the customer's needs, and suggest the most suitable smartphone model."

[0996] Thus, the system of the present invention enables staff to quickly propose the most suitable products that meet customer needs. This leads to improved customer satisfaction and more efficient customer service.

[0997] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0998] Step 1:

[0999] The device acquires conversations with the user in real time. The voice input device (microphone) collects the audio signal and converts it into a digital format. The input is the user's spoken voice, and the output is audio data in digital format.

[1000] Step 2:

[1001] The terminal encrypts the acquired digital audio data using an advanced encryption algorithm (e.g., AES-256) and sends it to the server using a secure communication protocol (e.g., HTTPS). The input is digital audio data, and the output is encrypted audio data.

[1002] Step 3:

[1003] The server receives encrypted audio data and first decrypts it. The input is the encrypted audio data, and the output is the decrypted audio data. Specifically, the AES-256 decryption algorithm is used.

[1004] Step 4:

[1005] The server inputs the decoded audio data into a speech recognition system (e.g., Google Cloud Speech-to-Text API) and converts the audio data into text data. The input is decoded audio data, and the output is text data. Specifically, it performs API calls.

[1006] Step 5:

[1007] The server analyzes the obtained text data using natural language processing techniques (e.g., spaCy, NLTK) and extracts keywords that indicate customer needs from the text data. The input is text data, and the output is the extracted keywords. Specifically, it performs text analysis and keyword extraction.

[1008] Step 6:

[1009] The server searches its internal database based on the extracted keywords. From the search results, it selects the optimal product using a predetermined evaluation algorithm (e.g., multi-criteria evaluation method). The input is the extracted keywords, and the output is information on the optimal product. Specifically, it performs optimization using database queries and evaluation algorithms.

[1010] Step 7:

[1011] The server converts the information of the selected products back into audio data using speech synthesis technology (e.g., Amazon Polly). The input is the optimal product information, and the output is audio data. Specifically, it makes API calls to convert text to speech.

[1012] Step 8:

[1013] The server sends encrypted audio data to the terminal. Input and output are the same as above, and encryption is performed again during the transmission process. Secure protocols such as HTTPS are used for communication.

[1014] Step 9:

[1015] The device decrypts the received encrypted audio data and plays it back using the smartphone's speaker. The input is the encrypted audio data, and the output is the played audio information. Specifically, it performs decryption and audio playback.

[1016] Step 10:

[1017] The terminal switches between smartphone suggestion mode and normal intercom mode via operation at the request of the staff. Specifically, mode switching is performed by pressing a physical button or touching the screen. Input is the staff's operation, and output is the mode change.

[1018] (Application Example 1)

[1019] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1020] Traditional brick-and-mortar stores face the challenge of staff being unable to quickly provide customers with appropriate product information. In particular, staff with limited product knowledge struggle to recommend the most suitable products for customer needs, leading to decreased customer satisfaction. Furthermore, the lack of systems capable of analyzing real-time conversational data and providing accurate product information hindered efficient service delivery.

[1021] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1022] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and extracting keywords, means for selecting the optimal product based on the extracted keywords, and means for using a generative AI model when converting product information into voice data. This makes it possible to analyze customer interactions in real time and provide optimal product information in voice.

[1023] A "terminal" is a device that acquires conversations with users as audio data and transmits it to a server.

[1024] "Audio data" refers to data obtained by converting the audio signal of a conversation with a user into a digital format.

[1025] A "server" is a computer system that receives audio data, converts it into text data, extracts keywords, and generates optimal product information based on that data.

[1026] "Text data" refers to written data converted from audio data by the server.

[1027] "Keywords" are important words or phrases extracted from text data that indicate customer needs or requests.

[1028] "Product information" refers to detailed information about the products selected by the server based on keywords.

[1029] A "generative AI model" refers to an algorithm or model that uses artificial intelligence to generate data, and in this invention, it is used specifically for generating audio data.

[1030] "Playback" refers to the operation in which a device processes the received audio data and outputs it as audio.

[1031] "Mode switching" refers to a function that allows a terminal to select different operating modes, and in this invention, it mainly involves switching between a product suggestion mode and a normal mode.

[1032] "Encryption" is the process by which audio data is digitally transformed to protect it from unauthorized access.

[1033] "Decryption" is the process of restoring encrypted data to its original state so that it can be interpreted as correct information.

[1034] "Natural language processing technology" is artificial intelligence technology that analyzes text data to understand its meaning and structure.

[1035] A "recommendation system" is an algorithm or software that selects and recommends the most suitable products or services based on user requirements.

[1036] This invention is a system for streamlining customer service in physical stores and improving customer satisfaction, providing product information using terminals, servers, and generative AI models. This system can analyze customer conversations using voice data and provide optimal product information.

[1037] System Configuration

[1038] terminal

[1039] The terminal is a device used by store staff to converse with customers. The terminal has the function to acquire voice data in real time, which is then encrypted and sent to a server. Furthermore, the terminal plays back the voice data received from the server, providing product information to the staff. The terminal also has a mode switching function, allowing it to switch between product suggestion mode and normal mode as needed.

[1040] server

[1041] The server receives audio data transmitted from the terminal and performs multiple analysis processes. First, it converts the audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing (NLP) technology to extract keywords that indicate customer needs. Based on the extracted keywords, the server searches its internal database and selects the most suitable product using its recommendation system. In this process, it uses a generative AI model to convert product information back into audio data.

[1042] Specific processing steps

[1043] The server uses standard cloud services such as "Google Cloud Speech-to-Text" for speech recognition to convert speech data into text data. Next, it uses software such as "Google Cloud Natural Language API" or "spaCy" for natural language processing to extract keywords from the text data. Then, based on a recommendation system, it searches the database for the best products that match the keywords and uses a generative AI model (e.g., "GPT-4") to convert the product information into speech data.

[1044] Specific examples of actions

[1045] 1. A customer enters the store and begins a conversation with a staff member.

[1046] Customer: "I'm looking for a smartphone with a great camera and long battery life."

[1047] 2. The device captures this conversation as audio data, encrypts it, and sends it to the server.

[1048] 3. The server converts the audio data into text data and uses NLP technology to extract keywords such as "camera performance" and "long battery life."

[1049] 4. The server searches the database based on keywords and selects the most suitable product. For example, "Model X" might be selected as the appropriate smartphone.

[1050] 5. The server converts the selected product information using the generation AI model into voice data and sends it to the terminal.

[1051] 6. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1052] Example of a prompt

[1053] Customer: "I'm looking for a smartphone with a great camera and long battery life."

[1054] System: "The Model X features a high-performance camera and a large-capacity battery."

[1055] In this way, store staff can respond quickly to customer needs.

[1056] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1057] Step 1:

[1058] The device acquires conversations with users as audio data. Specifically, store staff use a device (e.g., a smartphone) to interact with customers, and the device's microphone collects audio data during this interaction. The input is the conversation between the customer and the staff, and the output is the acquired audio data.

[1059] Step 2:

[1060] The terminal encrypts the acquired audio data and sends it to the server. Specifically, the terminal secures the audio data using an encryption algorithm such as AES and sends it to the server using the HTTPS protocol. The input is the acquired audio data, and the output is the encrypted audio data.

[1061] Step 3:

[1062] The server receives and decrypts encrypted audio data. Specifically, the server applies a decryption algorithm upon receipt to generate the original audio data. This ensures secure data transfer. The input is encrypted audio data, and the output is decrypted audio data.

[1063] Step 4:

[1064] The server converts the received audio data into text data using speech recognition technology. Specifically, it uses the "Google Cloud Speech-to-Text" API to analyze the audio data and transcribe it. The input is the decoded audio data, and the output is the corresponding text data.

[1065] Step 5:

[1066] The server analyzes text data using natural language processing (NLP) techniques to extract keywords. Specifically, it uses NLP software such as "Google Cloud Natural Language API" and "spaCy" to extract key keywords that indicate customer needs from the text data. The input is text data, and the output is the extracted keywords.

[1067] Step 6:

[1068] The server searches its internal database based on the extracted keywords and selects the most suitable product. Specifically, it uses a recommendation system (e.g., a filtering algorithm) to search for product information within the database and select the product that best matches the keywords. The input is the extracted keywords, and the output is the selected product information.

[1069] Step 7:

[1070] The server converts selected product information into audio data using a generative AI model. Specifically, it uses a generative AI model such as "GPT-4" to convert the product information into text, and then uses the "Google Cloud Text-to-Speech" API to convert it into audio data. The input is selected product information, and the output is audio data.

[1071] Step 8:

[1072] The server sends the generated audio data to the terminal. Specifically, it re-encrypts the audio data and sends it to the terminal via the HTTPS protocol. The input is the generated audio data, and the output is the encrypted audio data.

[1073] Step 9:

[1074] The terminal decrypts and plays back the received audio data. Specifically, it decrypts the received encrypted data and provides product information to staff via audio using the terminal's speaker. The input is the received encrypted audio data, and the output is the played audio information.

[1075] This series of processing steps enables store staff to provide customers with product information quickly and accurately in response to their needs.

[1076] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1077] The system of this invention aims to analyze conversations with customers and provide optimal product information. This system includes terminals, servers, an emotion engine, and multiple technical means that utilize voice and text data.

[1078] Specifically, the system operates as follows:

[1079] Terminal operation

[1080] The terminal captures audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and transmitted to the server in an encrypted state.

[1081] Server Processing

[1082] The server first converts the received audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing technology to extract keywords that indicate customer needs.

[1083] How the emotion engine works

[1084] Furthermore, the voice data is analyzed by an emotion engine to recognize the customer's emotional state (e.g., satisfaction, dissatisfaction, excitement, questioning, etc.). Based on this emotional information, the emotion engine adjusts the importance and priority of keywords in product recommendations.

[1085] Selection of the optimal product

[1086] The server searches its internal database by combining extracted keywords and sentiment information, and selects the optimal smartphone model using a predetermined algorithm. For example, conditions such as "camera performance," "battery life," and "price range" may be extracted as keywords. If the sentiment engine recognizes that the customer is excited, it may make suggestions, including bold promotions.

[1087] Product Information

[1088] The information from the selected smartphones is converted back into audio data using speech synthesis technology on the server, encrypted, and sent to the terminal. The terminal plays this received audio data and presents the product information to the staff verbally.

[1089] The emotion engine adjusts its tone according to the user's emotions; if the user appears anxious, it adjusts to provide information in a calmer voice. This is expected to further enhance user satisfaction.

[1090] Mode switching

[1091] The terminal can switch modes with the press of a button according to the staff's request. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[1092] Specific example

[1093] The following are some specific use cases.

[1094] 1. The user (customer) enters the store and begins a conversation with the staff.

[1095] User: "I'm looking for a smartphone with a great camera and long battery life."

[1096] 2. The device captures this conversation as audio data and sends it to the server.

[1097] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[1098] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[1099] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[1100] 6. The server converts the selection results into audio data and sends it to the terminal. At this time, the emotion engine applies a tone that corresponds to the user's excitement level.

[1101] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1102] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[1103] In this way, staff can quickly and accurately suggest smartphones that match the customer's needs and feelings. This system makes it possible to provide high-quality customer service without relying on the staff's knowledge.

[1104] ---

[1105] The above describes specific embodiments for carrying out the present invention.

[1106] The following describes the processing flow.

[1107] Step 1:

[1108] The terminal begins voice input when a staff member starts a conversation with a customer. The conversation is captured in real time as audio data in digital format.

[1109] Step 2:

[1110] The device encrypts the acquired audio data and securely transmits it to the server.

[1111] Step 3:

[1112] The server decrypts the received encrypted audio data. Then, it uses speech recognition technology to convert the audio data into text data.

[1113] Step 4:

[1114] The server analyzes the converted text data and uses natural language processing techniques to extract important keywords. For example, keywords such as "camera performance," "battery life," and "price" may be extracted.

[1115] Step 5:

[1116] The server simultaneously passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine recognizes the customer's emotions, such as excitement, satisfaction, dissatisfaction, or questioning, based on their voice tone and speaking style.

[1117] Step 6:

[1118] The server searches its internal database based on the extracted keywords and sentiment information obtained from the sentiment engine, and lists the relevant smartphone models.

[1119] Step 7:

[1120] The server selects the optimal smartphone model from the listed models using an evaluation algorithm that combines emotional information. For example, if the customer appears anxious, a model emphasizing ease of use will be selected.

[1121] Step 8:

[1122] The server generates information about the selected smartphone model as text data. For example, it might include information such as, "Model X features a high-performance camera and a large-capacity battery."

[1123] Step 9:

[1124] The server converts the generated text data into speech data using speech synthesis technology. During this process, an emotion engine adjusts the tone of the speech. For example, if the user is excited, the speech will be generated in an energetic tone; if they are calm, it will be generated in a gentle tone.

[1125] Step 10:

[1126] The server encrypts the audio data and sends it to the terminal.

[1127] Step 11:

[1128] The device decrypts the received encrypted audio data and prepares it for playback.

[1129] Step 12:

[1130] The device transmits the played audio data to staff via an intercom. This allows staff to make real-time recommendations for the most suitable smartphone for each customer.

[1131] Step 13:

[1132] The terminal allows staff to easily switch between smartphone suggestion mode and normal intercom mode by operating a button on the intercom.

[1133] Specific example:

[1134] 1. The user (customer) enters the store and begins a conversation with the staff.

[1135] User: "I'm looking for a smartphone with a great camera and long battery life."

[1136] 2. The device captures this conversation as audio data and sends it to the server.

[1137] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[1138] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[1139] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[1140] 6. The server converts the selection results into audio data, applies a tone corresponding to the user's excitement level using the emotion engine, and sends it to the terminal.

[1141] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1142] 8. The terminal can be switched from smartphone suggestion mode to normal intercom mode by the staff operating a button on the intercom.

[1143] In this way, staff can quickly and accurately suggest smartphones that match the customer's needs and feelings. This system makes it possible to provide high-quality customer service without relying on the staff's knowledge.

[1144] (Example 2)

[1145] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1146] Traditional customer service systems rely heavily on staff knowledge for product recommendations and responses, making it difficult to provide high-quality service that responds promptly to customer needs and emotions. Furthermore, manually converting voice data to text and performing sentiment analysis is time-consuming and inaccurate. This can lead to decreased customer satisfaction and negatively impact store sales. Additionally, insufficient security measures for transmitting and receiving voice data create a risk of information leakage.

[1147] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1148] In this invention, the server includes means for converting voice data into text data using voice recognition technology, means for analyzing the text data and extracting keywords using natural language processing technology, and means for analyzing the voice data using an emotion analysis engine to recognize the customer's emotional state. This makes it possible to quickly respond to customer needs and emotions and provide high-quality, secure customer service.

[1149] A "terminal" is a device that captures conversations between staff and customers and transmits the audio data to a server.

[1150] A "server" is a central processing unit that receives and analyzes audio data and provides optimal product information.

[1151] "Voice data" refers to digital signals acquired by the terminal, including the content of conversations between staff and customers.

[1152] "Encryption" refers to the process of transforming and protecting audio data for security purposes.

[1153] "Decryption" is the process of restoring encrypted audio data to its original state.

[1154] "Speech recognition technology" is a technology that converts speech data into text data.

[1155] "Text data" refers to data that represents conversation content as written information, converted using speech recognition technology.

[1156] "Natural language processing technology" is a general term for technologies that analyze text data and extract keywords.

[1157] "Keywords" are important terms that indicate customer needs.

[1158] An "emotion engine" is a system that analyzes voice data to recognize the customer's emotional state.

[1159] "Emotional state" refers to the mental state a customer expresses during a conversation, such as "satisfaction," "dissatisfaction," "excitement," or "question."

[1160] A "database" is a data storage system used to store product information and other necessary information.

[1161] An "algorithm" refers to a set of calculation and processing steps used to select the optimal product based on extracted keywords and emotional information.

[1162] "Speech synthesis technology" is a technology that converts text data into speech data.

[1163] "Playback" is the process by which a device outputs received audio data as audio.

[1164] "Mode" refers to the operating state of the device, and includes, for example, smartphone suggestion mode and normal intercom mode.

[1165] The system of the present invention is built to provide optimal product information through interaction with users (customers). This system includes terminals, servers, an emotion engine, and multiple technical means that utilize voice and text data.

[1166] Specifically, the system operates using the following hardware and software.

[1167] Terminal operation

[1168] The terminal captures audio in real time as store staff converse with customers. The hardware used includes a high-quality microphone with noise-canceling capabilities. The captured audio data is converted to a digital format using digital signal processing technology and transmitted to the server with enhanced security using AES-256 encryption technology.

[1169] Server Processing

[1170] The server receives encrypted audio data sent from the terminal and first decrypts it. Then, it converts the audio data into text data using the Google Speech-to-Text API. This conversion process removes noise and maintains quality. Next, it analyzes the text data using a natural language processing library (e.g., spaCy or NLTK) to extract keywords that indicate customer needs. The extracted keywords specifically describe the product features that the customer is looking for.

[1171] How the emotion engine works

[1172] The server has a built-in emotion engine (e.g., IBM Watson Tone Analyzer) that analyzes the customer's emotional state based on voice data. The emotion engine analyzes the customer's voice tone, speed, pitch, etc., and determines emotional states such as "satisfied," "dissatisfied," "excited," and "questionable." This emotional information is used to adjust the importance and priority of keywords in the product proposal process described later.

[1173] Selection of the optimal product

[1174] The server searches its internal database based on extracted keywords and sentiment information, and selects the optimal product using a pre-configured algorithm (e.g., weighting calculations and evaluation functions). For example, if a customer prioritizes "camera performance" and "battery life," and the sentiment engine detects excitement, a product proposal, including a bold promotion, will be made.

[1175] Product Information

[1176] Information on selected products is converted back into audio data by the server using the Google Text-to-Speech API, and an emotion engine sets a tone that matches the customer's emotions. The re-encrypted audio data is sent to the terminal. The terminal decrypts the received audio data and plays it back through the staff member's earphones or speaker, providing the product information. Text information is also displayed on the terminal's screen.

[1177] Mode switching

[1178] The terminal features an interface that can be operated according to the staff's requests, and allows for easy switching between smartphone suggestion mode and normal intercom mode.

[1179] Specific example

[1180] The following are some specific use cases.

[1181] 1. The user (customer) enters the store and begins a conversation with the staff.

[1182] User: "I'm looking for a smartphone with a great camera and long battery life."

[1183] 2. The device captures this conversation as audio data and sends it to the server. This is done using the built-in microphone to capture high-quality audio data, which is then encrypted.

[1184] 3. The server converts the audio data into text data using the Google Speech-to-Text API and performs noise filtering. Next, it uses natural language processing technology (e.g., spaCy) to extract keywords such as "camera performance" and "long battery life."

[1185] 4. An emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes that the user is in an excited state.

[1186] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm (e.g., a weighting algorithm). For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[1187] 6. The server converts the selection results into audio data using the Google Text-to-Speech API and sends it to the device. At this time, the emotion engine applies a tone that corresponds to the user's level of excitement.

[1188] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1189] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[1190] Examples of prompts for generative AI models

[1191] The following are examples of prompts for a generative AI model.

[1192] User: "I'm looking for a smartphone with a great camera and long battery life."

[1193] This system allows staff to quickly and accurately recommend smartphones that match the customer's needs and feelings. It also enables high-quality customer service without relying on the staff's knowledge.

[1194] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1195] Program processing flow

[1196] Step 1: The device acquires the audio.

[1197] The terminal captures conversations between staff and users (customers) in real time using its built-in microphone. The audio data is converted into a digital signal and encrypted at high speed.

[1198] Input: Audio of a conversation between staff and customer

[1199] Processing: Digital signal conversion and encryption of audio (using AES-256 encryption technology)

[1200] Output: Encrypted digital audio data

[1201] Step 2: The device sends the audio data to the server.

[1202] The device sends encrypted voice data to the server. This communication is conducted using a secure protocol.

[1203] Input: Encrypted digital audio data

[1204] Processing: Data transmission (using a secure communication protocol)

[1205] Output: Data arrives on the server.

[1206] Step 3: The server decrypts the audio data.

[1207] The server receives encrypted audio data sent from the terminal and decrypts it using AES-256 decryption technology.

[1208] Input: Encrypted digital audio data

[1209] Processing: Decoding of audio data (using AES-256 decoding technology)

[1210] Output: Decoded digital audio data

[1211] Step 4: The server converts the audio to text.

[1212] The server uses the Google Speech-to-Text API to convert the audio data into text data.

[1213] Input: Decoded digital audio data

[1214] Processing: Speech recognition (using Google Speech-to-Text API)

[1215] Output: Text data

[1216] Step 5: The server extracts the keywords.

[1217] The server uses natural language processing tools such as spaCy and NLTK to extract important keywords from text data.

[1218] Input: Text data

[1219] Processing: Natural language processing (keyword extraction, TF-IDF calculation, etc.)

[1220] Output: Extracted keywords

[1221] Step 6: The emotion engine analyzes emotions.

[1222] The server uses an emotion engine to analyze voice data and recognize the user's emotional state.

[1223] Input: Decoded digital audio data

[1224] Processing: Sentiment analysis (analyzes voice tone, word choice, speed, etc.)

[1225] Output: Emotional information (satisfied, dissatisfied, excited, questioned, etc.)

[1226] Step 7: Select the optimal product for the server.

[1227] The server searches its internal database based on the extracted keywords and sentiment information to select the most suitable product.

[1228] Input: Extracted keywords, sentiment information

[1229] Processing: Database search and evaluation algorithms (weighting calculation, similarity calculation, etc.)

[1230] Output: Optimal product information

[1231] Step 8: The server generates the audio data.

[1232] The server uses the Google Text-to-Speech API to convert selected product information into audio data. The tone is then adjusted by an emotion engine.

[1233] Input: Optimal product information

[1234] Processing: Text-to-speech (using Google Text-to-Speech API)

[1235] Output: Audio data

[1236] Step 9: The server sends encrypted audio data to the device.

[1237] The server re-encrypts the generated audio data and sends it to the terminal.

[1238] Input: Audio data

[1239] Processing: Encryption of audio data (using AES-256 encryption technology) and transmission.

[1240] Output: Encrypted audio data

[1241] Step 10: The device plays the audio data.

[1242] The terminal decodes the received audio data and plays it back through earphones or speakers worn by the staff. It also displays text information on its screen.

[1243] Input: Encrypted audio data

[1244] Processing: Decoding and playback of audio data, displaying text on the screen.

[1245] Output: Played audio, displayed text information

[1246] Step 11: The device switches modes.

[1247] The terminal's mode can be switched by staff operation. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[1248] Input: Staff operation (button press)

[1249] Processing: Mode switching

[1250] Output: New mode setting

[1251] Example of a prompt

[1252] The following are examples of prompts for a generative AI model.

[1253] User: "I'm looking for a smartphone with a great camera and long battery life."

[1254] (Application Example 2)

[1255] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1256] Conventional customer service systems often rely on subjective judgments from staff when recommending products, failing to adequately reflect customers' specific needs and emotional states. Furthermore, providing timely and appropriate product information is difficult, resulting in a lack of effective means to improve customer satisfaction. This invention aims to solve these problems by providing a system that delivers optimal product information in real time, based on customer needs and emotions.

[1257] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's emotional state using an emotion engine and adjusting the tone of product suggestions based on the emotional state; means for generating the tone of product suggestions as prompt sentences using a generation AI model; and means for extracting keywords using natural language processing technology and selecting the optimal product using a recommendation system. This makes it possible to make product suggestions in an appropriate tone that takes into account the customer's emotional state, and is expected to improve customer satisfaction.

[1258] A "terminal" is a device that acquires conversations with a user as audio data, sends that data to a server, and plays back information from the server.

[1259] A "server" is a central processing unit that converts audio data into text data, analyzes that text data to extract keywords, analyzes the user's emotional state using an emotion engine, selects the optimal product, and transmits that information to the terminal.

[1260] An "emotion engine" is an algorithm or software that can analyze a user's voice data and recognize their emotional state.

[1261] "Keywords" are important words and phrases extracted from user conversations that are useful when selecting or proposing products.

[1262] "Natural language processing technology" is a technique for analyzing text data and extracting meaningful keywords from it.

[1263] A "recommendation system" is an algorithm or software that searches a database based on extracted keywords and selects the most suitable product.

[1264] A "generative AI model" is a machine learning-based artificial intelligence model used to generate the tone of product proposals.

[1265] A "prompt sentence" is an input sentence used to generate an AI model that includes the tone of a product proposal.

[1266] "Means for switching modes" refers to methods for changing the operating mode of a device depending on its intended use, such as switching between smartphone suggestion mode and normal intercom mode.

[1267] A system for specifically implementing the present invention is described below. This system acquires conversations between customers and staff in real time and provides optimal product information by analyzing the content of the conversation and the customer's emotional state.

[1268] Terminal operation

[1269] The terminal first captures the audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and sent to the server in an encrypted state. Standard protocols such as SSL / TLS are used for this encryption.

[1270] Server Processing

[1271] The server first converts the received audio data into text data using speech recognition technology. Google's speech recognition API is used here. Next, this text data is analyzed using natural language processing technology to extract keywords that indicate customer needs. The Python nltk library is used for natural language processing.

[1272] Furthermore, the server uses an emotion engine to analyze the voice data and recognize the customer's emotional state (e.g., satisfaction, dissatisfaction, excitement, questioning, etc.). The emotion engine uses the Hugging Face transformers library.

[1273] Selection of the optimal product

[1274] The server searches its internal database based on extracted keywords and sentiment information, and selects the optimal product using a predetermined algorithm. During this process, the database search and recommendation system work in conjunction. The recommendation system utilizes the scikit-learn library.

[1275] The selected product information is then adjusted by an emotion engine to match the customer's emotional tone. For example, if the customer is excited, the suggestions will be presented in an energetic tone. These suggestions are generated as prompts based on a generative AI model.

[1276] Product Information

[1277] The information on the selected products is converted back into audio data by the server using speech synthesis technology. The pyttsx3 library is used for this speech synthesis. Next, the audio data is encrypted again and sent to the terminal. The terminal plays the received audio data and presents the product information to the staff verbally.

[1278] Mode switching

[1279] The device has a function that allows staff to switch modes with the press of a button, according to their request. This makes it easy to switch between smartphone suggestion mode and normal intercom mode.

[1280] Specific example

[1281] The following are some specific use cases.

[1282] 1. The user (customer) enters the store and begins a conversation with the staff.

[1283] User: "I'm looking for a smartphone with a great camera and long battery life."

[1284] 2. The device captures this conversation as audio data and sends it to the server.

[1285] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[1286] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[1287] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[1288] 6. The server converts the selection results into audio data and sends it to the terminal. At this time, the emotion engine applies a tone that corresponds to the user's excitement level.

[1289] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1290] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[1291] Example of a prompt

[1292] For example, if a customer says, "I'm looking for a smartphone with a great camera and long battery life," the following prompt will be generated.

[1293] A customer says, "I'm looking for a smartphone with excellent camera performance and long battery life." Please recommend the best product for them. Use sentiment analysis to ensure your recommendation is delivered in an appropriate tone.

[1294] This system allows staff to quickly and accurately suggest products that match the customer's needs and feelings. This is expected to improve customer satisfaction.

[1295] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1296] Step 1:

[1297] The terminal acquires the conversation with the user as audio data in real time. The input is the audio data of the conversation between the user and the staff, and the output is the audio data converted into a digital signal.

[1298] Step 2:

[1299] The terminal encrypts the acquired audio data and sends it to the server. The input is audio data converted into a digital signal, and the output is encrypted audio data. Standard protocols such as SSL / TLS are used for encryption.

[1300] Step 3:

[1301] The server converts received audio data into text data using speech recognition technology. The input is encrypted audio data, and the output is text data. Google's speech recognition API is used for speech recognition.

[1302] Step 4:

[1303] The server analyzes text data using natural language processing techniques and extracts keywords. The input is text data obtained through speech recognition technology, and the output is the analyzed keywords. The NLTK library is used for natural language processing.

[1304] Step 5:

[1305] The server analyzes voice data using an emotion engine to recognize the user's emotional state. The input is voice data acquired in real time, and the output is the recognized emotional state. The Hugging Face transformers library is used as the emotion engine.

[1306] Step 6:

[1307] The server combines extracted keywords and sentiment information to search an internal database and selects the optimal product using an evaluation algorithm. The input is keywords and sentiment information, and the output is information on the selected product. The recommendation system uses the scikit-learn library.

[1308] Step 7:

[1309] The server uses a generative AI model to adjust the tone of product suggestions and generate prompt sentences. The input is selected product information and sentiment information, and the output is the prompt sentence. A pre-trained language model is used for the generative AI model.

[1310] Step 8:

[1311] The server converts product information back into audio data, encrypts it, and sends it to the terminal. The input is a prompt message, and the output is encrypted audio data. The pyttsx3 library is used for speech synthesis.

[1312] Step 9:

[1313] The terminal decrypts and plays back the received audio data. The input is encrypted audio data, and the output is the decrypted and played audio information. This audio information is presented to staff as a product suggestion via the terminal.

[1314] Step 10:

[1315] Staff members can switch between smartphone suggestion mode and normal intercom mode by operating a button on the device. The input is the button operation signal, and the output is the mode switching status of the device.

[1316] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1317] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1318] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1319] [Fourth Embodiment]

[1320] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1321] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1322] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1323] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1324] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1325] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1326] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1327] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1328] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1329] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1330] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1331] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1332] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1333] The system of this invention aims to analyze conversations with customers and provide optimal product information. This system includes terminals, servers, and multiple technical means that utilize voice and text data.

[1334] Specifically, the system operates as follows:

[1335] Terminal operation

[1336] The terminal captures audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and transmitted to the server in an encrypted state.

[1337] Server Processing

[1338] The server first converts the received audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing technology to extract keywords that indicate customer needs.

[1339] Based on the extracted keywords, the server searches its internal database and selects the optimal smartphone model using a predetermined algorithm. For example, conditions such as "camera performance," "battery life," and "price range" may be extracted as keywords. In this case, smartphones with these characteristics will be selected.

[1340] Product Information

[1341] The information from the selected smartphones is converted back into audio data using speech synthesis technology on the server, encrypted, and sent to the terminal. The terminal plays this received audio data and presents the product information to the staff verbally.

[1342] Mode switching

[1343] The terminal can switch modes with the press of a button according to the staff's request. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[1344] Specific example

[1345] The following are some specific use cases.

[1346] 1. The user (customer) enters the store and begins a conversation with the staff.

[1347] User: "I'm looking for a smartphone with a great camera and long battery life."

[1348] 2. The device captures this conversation as audio data and sends it to the server.

[1349] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[1350] 4. The server searches the database for multiple smartphone models based on keywords and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[1351] 5. The server converts the selection results into audio data and sends it to the terminal.

[1352] 6. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1353] In this way, staff can quickly suggest smartphones that meet customer needs. This system makes it possible to provide high-quality customer service without relying on the staff's knowledge.

[1354] ---

[1355] The above describes specific embodiments for carrying out the present invention.

[1356] The following describes the processing flow.

[1357] Step 1:

[1358] The terminal begins voice input when a staff member starts a conversation with a customer. The conversation is captured in real time as audio data in digital format.

[1359] Step 2:

[1360] The device encrypts the acquired audio data and securely transmits it to the server.

[1361] Step 3:

[1362] The server decrypts the received encrypted audio data. Then, it uses speech recognition technology to convert the audio data into text data.

[1363] Step 4:

[1364] The server analyzes the converted text data and uses natural language processing techniques to extract important keywords. For example, keywords such as "camera performance," "battery life," and "price" may be extracted.

[1365] Step 5:

[1366] The server searches its internal database based on the extracted keywords and lists the relevant smartphone models.

[1367] Step 6:

[1368] The server uses an evaluation algorithm to select the optimal smartphone model from the listed models.

[1369] Step 7:

[1370] The server generates text data describing the features of the selected smartphone model. For example, it might include information such as, "Model X features a high-performance camera and a large-capacity battery."

[1371] Step 8:

[1372] The server converts the generated text data into speech data using speech synthesis technology. Then, it encrypts the speech data and sends it to the terminal.

[1373] Step 9:

[1374] The device decrypts the received encrypted audio data and prepares it for playback as audio.

[1375] Step 10:

[1376] The device transmits the played audio data to staff via an intercom. This allows staff to make real-time recommendations for the most suitable smartphone for each customer.

[1377] Step 11:

[1378] The terminal allows staff to easily switch between smartphone suggestion mode and normal intercom mode by operating a button on the intercom.

[1379] (Example 1)

[1380] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1381] In modern retail, responding promptly to customer needs is crucial, but not all staff possess extensive product knowledge. This can lead to delays in recommending the best product or even provide incorrect information. Furthermore, a mismatch between staff information and customer requests can reduce customer satisfaction. Additionally, there's a lack of systems to efficiently process useful information gathered during conversations and provide optimal product recommendations.

[1382] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1383] In this invention, the server includes means for converting speech into text data using a speech recognition system, means for extracting keywords from the text data using natural language processing technology, and means for searching an internal database and selecting the optimal product using an evaluation algorithm. This enables the system to accurately grasp customer requirements and make quick and accurate product suggestions without relying on the knowledge of staff.

[1384] A "terminal" is a device that acquires conversations with users as audio data and transmits it to a server.

[1385] A "server" is a computer system that receives audio data sent from a terminal, converts it to text data, analyzes it, selects products, and converts it back to audio data.

[1386] "Audio data" refers to data obtained by converting the voices of users or staff into digital signals.

[1387] "Text data" refers to character information converted from speech data using speech recognition technology.

[1388] "Keywords" are important words or phrases extracted from text data that indicate customer needs.

[1389] A "product" is a product that the server proposes based on customer needs, and in this example, it is a smartphone.

[1390] A "speech recognition system" is a system that possesses the technology to convert speech data into text data.

[1391] "Natural language processing technology" is a technique that analyzes text data and extracts keywords.

[1392] "Speech synthesis technology" is a technology that converts text data into speech data.

[1393] An "internal database" is a data store containing product information stored within the server.

[1394] An "evaluation algorithm" is a computational method for selecting the optimal product based on extracted keywords.

[1395] "Encryption" is the process of transforming data for security purposes.

[1396] "Decryption" is the process of restoring encrypted data to its original form.

[1397] The system of this invention aims to analyze conversations with customers and provide optimal product information. This system consists of a terminal, a server, and multiple technical means such as speech recognition, natural language processing, database search, and speech synthesis technology.

[1398] Terminal operation

[1399] The terminal uses a voice input device (microphone) to capture conversations with the user in real time. The captured voice data is converted into a digital signal and then encrypted. An advanced algorithm such as AES-256 is used for encryption, and the data is sent to the server via a secure communication protocol (e.g., HTTPS).

[1400] Server Processing

[1401] Speech recognition and text conversion

[1402] The server receives the encrypted audio data and first decrypts it. The decrypted audio data is then input into a speech recognition system (e.g., Google Cloud Speech-to-Text API) to convert the audio data into text data.

[1403] Text data analysis

[1404] The server analyzes the obtained text data using natural language processing techniques (e.g., spaCy, NLTK) to extract keywords that indicate customer needs. This keyword extraction clarifies the user's requirements.

[1405] Product selection

[1406] Based on the extracted keywords, the server searches its internal database. This database contains detailed product information. The search results are analyzed using a predetermined evaluation algorithm (e.g., multi-criteria evaluation method) to select the optimal product.

[1407] Product Information

[1408] Reconversion to audio data

[1409] The information on the selected products is converted back into audio data by the server using speech synthesis technology (e.g., Amazon Polly). This audio data is then encrypted again and sent to the device.

[1410] Playback of audio data

[1411] The device decrypts the received encrypted audio data and plays it back using the smartphone's speaker. Based on this audio, staff can then provide customers with the most relevant product information.

[1412] Mode switching

[1413] The terminal can be switched between smartphone suggestion mode and normal intercom mode by staff depending on the usage situation. This function improves the convenience of customer service.

[1414] Specific example

[1415] The following are some specific use cases:

[1416] 1. The user (customer) enters the store and begins a conversation with the staff.

[1417] User: "I'm looking for a smartphone with a great camera and long battery life."

[1418] 2. The device captures this conversation using its microphone, encrypts the audio data, and sends it to the server.

[1419] 3. The server receives the audio data and converts it into text data using a speech recognition system.

[1420] Example translation: "I'm looking for a smartphone with a great camera and long battery life."

[1421] 4. The server extracts the keywords "camera performance" and "long battery life" from the retrieved text.

[1422] 5. The server searches the database and determines that "Model X" is the optimal model based on its evaluation algorithm.

[1423] Example of selection reason: Model X has a high-performance camera and a large-capacity battery.

[1424] 6. The server converts the information from Model X back into audio data, encrypts it, and transmits it.

[1425] 7. The device decodes the audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1426] Examples of prompt statements

[1427] The following are examples of prompts to input into the generated AI model:

[1428] "Convert this audio data into text data, extract keywords that align with the customer's needs, and suggest the most suitable smartphone model."

[1429] Thus, the system of the present invention enables staff to quickly propose the most suitable products that meet customer needs. This leads to improved customer satisfaction and more efficient customer service.

[1430] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1431] Step 1:

[1432] The device acquires conversations with the user in real time. The voice input device (microphone) collects the audio signal and converts it into a digital format. The input is the user's spoken voice, and the output is audio data in digital format.

[1433] Step 2:

[1434] The terminal encrypts the acquired digital audio data using an advanced encryption algorithm (e.g., AES-256) and sends it to the server using a secure communication protocol (e.g., HTTPS). The input is digital audio data, and the output is encrypted audio data.

[1435] Step 3:

[1436] The server receives encrypted audio data and first decrypts it. The input is the encrypted audio data, and the output is the decrypted audio data. Specifically, the AES-256 decryption algorithm is used.

[1437] Step 4:

[1438] The server inputs the decoded audio data into a speech recognition system (e.g., Google Cloud Speech-to-Text API) and converts the audio data into text data. The input is decoded audio data, and the output is text data. Specifically, it performs API calls.

[1439] Step 5:

[1440] The server analyzes the obtained text data using natural language processing techniques (e.g., spaCy, NLTK) and extracts keywords that indicate customer needs from the text data. The input is text data, and the output is the extracted keywords. Specifically, it performs text analysis and keyword extraction.

[1441] Step 6:

[1442] The server searches its internal database based on the extracted keywords. From the search results, it selects the optimal product using a predetermined evaluation algorithm (e.g., multi-criteria evaluation method). The input is the extracted keywords, and the output is information on the optimal product. Specifically, it performs optimization using database queries and evaluation algorithms.

[1443] Step 7:

[1444] The server converts the information of the selected products back into audio data using speech synthesis technology (e.g., Amazon Polly). The input is the optimal product information, and the output is audio data. Specifically, it makes API calls to convert text to speech.

[1445] Step 8:

[1446] The server sends encrypted audio data to the terminal. Input and output are the same as above, and encryption is performed again during the transmission process. Secure protocols such as HTTPS are used for communication.

[1447] Step 9:

[1448] The device decrypts the received encrypted audio data and plays it back using the smartphone's speaker. The input is the encrypted audio data, and the output is the played audio information. Specifically, it performs decryption and audio playback.

[1449] Step 10:

[1450] The terminal switches between smartphone suggestion mode and normal intercom mode via operation at the request of the staff. Specifically, mode switching is performed by pressing a physical button or touching the screen. Input is the staff's operation, and output is the mode change.

[1451] (Application Example 1)

[1452] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1453] Traditional brick-and-mortar stores face the challenge of staff being unable to quickly provide customers with appropriate product information. In particular, staff with limited product knowledge struggle to recommend the most suitable products for customer needs, leading to decreased customer satisfaction. Furthermore, the lack of systems capable of analyzing real-time conversational data and providing accurate product information hindered efficient service delivery.

[1454] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1455] In this invention, the server includes means for converting voice data into text data, means for analyzing the text data and extracting keywords, means for selecting the optimal product based on the extracted keywords, and means for using a generative AI model when converting product information into voice data. This makes it possible to analyze customer interactions in real time and provide optimal product information in voice.

[1456] A "terminal" is a device that acquires conversations with users as audio data and transmits it to a server.

[1457] "Audio data" refers to data obtained by converting the audio signal of a conversation with a user into a digital format.

[1458] A "server" is a computer system that receives audio data, converts it into text data, extracts keywords, and generates optimal product information based on that data.

[1459] "Text data" refers to written data converted from audio data by the server.

[1460] "Keywords" are important words or phrases extracted from text data that indicate customer needs or requests.

[1461] "Product information" refers to detailed information about the products selected by the server based on keywords.

[1462] A "generative AI model" refers to an algorithm or model that uses artificial intelligence to generate data, and in this invention, it is used specifically for generating audio data.

[1463] "Playback" refers to the operation in which a device processes the received audio data and outputs it as audio.

[1464] "Mode switching" refers to a function that allows a terminal to select different operating modes, and in this invention, it mainly involves switching between a product suggestion mode and a normal mode.

[1465] "Encryption" is the process by which audio data is digitally transformed to protect it from unauthorized access.

[1466] "Decryption" is the process of restoring encrypted data to its original state so that it can be interpreted as correct information.

[1467] "Natural language processing technology" is artificial intelligence technology that analyzes text data to understand its meaning and structure.

[1468] A "recommendation system" is an algorithm or software that selects and recommends the most suitable products or services based on user requirements.

[1469] This invention is a system for streamlining customer service in physical stores and improving customer satisfaction, providing product information using terminals, servers, and generative AI models. This system can analyze customer conversations using voice data and provide optimal product information.

[1470] System Configuration

[1471] terminal

[1472] The terminal is a device used by store staff to converse with customers. The terminal has the function to acquire voice data in real time, which is then encrypted and sent to a server. Furthermore, the terminal plays back the voice data received from the server, providing product information to the staff. The terminal also has a mode switching function, allowing it to switch between product suggestion mode and normal mode as needed.

[1473] server

[1474] The server receives audio data transmitted from the terminal and performs multiple analysis processes. First, it converts the audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing (NLP) technology to extract keywords that indicate customer needs. Based on the extracted keywords, the server searches its internal database and selects the most suitable product using its recommendation system. In this process, it uses a generative AI model to convert product information back into audio data.

[1475] Specific processing steps

[1476] The server uses standard cloud services such as "Google Cloud Speech-to-Text" for speech recognition to convert speech data into text data. Next, it uses software such as "Google Cloud Natural Language API" or "spaCy" for natural language processing to extract keywords from the text data. Then, based on a recommendation system, it searches the database for the best products that match the keywords and uses a generative AI model (e.g., "GPT-4") to convert the product information into speech data.

[1477] Specific examples of actions

[1478] 1. A customer enters the store and begins a conversation with a staff member.

[1479] Customer: "I'm looking for a smartphone with a great camera and long battery life."

[1480] 2. The device captures this conversation as audio data, encrypts it, and sends it to the server.

[1481] 3. The server converts the audio data into text data and uses NLP technology to extract keywords such as "camera performance" and "long battery life."

[1482] 4. The server searches the database based on keywords and selects the most suitable product. For example, "Model X" might be selected as the appropriate smartphone.

[1483] 5. The server converts the selected product information using the generation AI model into voice data and sends it to the terminal.

[1484] 6. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1485] Example of a prompt

[1486] Customer: "I'm looking for a smartphone with a great camera and long battery life."

[1487] System: "The Model X features a high-performance camera and a large-capacity battery."

[1488] In this way, store staff can respond quickly to customer needs.

[1489] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1490] Step 1:

[1491] The device acquires conversations with users as audio data. Specifically, store staff use a device (e.g., a smartphone) to interact with customers, and the device's microphone collects audio data during this interaction. The input is the conversation between the customer and the staff, and the output is the acquired audio data.

[1492] Step 2:

[1493] The terminal encrypts the acquired audio data and sends it to the server. Specifically, the terminal secures the audio data using an encryption algorithm such as AES and sends it to the server using the HTTPS protocol. The input is the acquired audio data, and the output is the encrypted audio data.

[1494] Step 3:

[1495] The server receives and decrypts encrypted audio data. Specifically, the server applies a decryption algorithm upon receipt to generate the original audio data. This ensures secure data transfer. The input is encrypted audio data, and the output is decrypted audio data.

[1496] Step 4:

[1497] The server converts the received audio data into text data using speech recognition technology. Specifically, it uses the "Google Cloud Speech-to-Text" API to analyze the audio data and transcribe it. The input is the decoded audio data, and the output is the corresponding text data.

[1498] Step 5:

[1499] The server analyzes text data using natural language processing (NLP) techniques to extract keywords. Specifically, it uses NLP software such as "Google Cloud Natural Language API" and "spaCy" to extract key keywords that indicate customer needs from the text data. The input is text data, and the output is the extracted keywords.

[1500] Step 6:

[1501] The server searches its internal database based on the extracted keywords and selects the most suitable product. Specifically, it uses a recommendation system (e.g., a filtering algorithm) to search for product information within the database and select the product that best matches the keywords. The input is the extracted keywords, and the output is the selected product information.

[1502] Step 7:

[1503] The server converts selected product information into audio data using a generative AI model. Specifically, it uses a generative AI model such as "GPT-4" to convert the product information into text, and then uses the "Google Cloud Text-to-Speech" API to convert it into audio data. The input is selected product information, and the output is audio data.

[1504] Step 8:

[1505] The server sends the generated audio data to the terminal. Specifically, it re-encrypts the audio data and sends it to the terminal via the HTTPS protocol. The input is the generated audio data, and the output is the encrypted audio data.

[1506] Step 9:

[1507] The terminal decrypts and plays back the received audio data. Specifically, it decrypts the received encrypted data and provides product information to staff via audio using the terminal's speaker. The input is the received encrypted audio data, and the output is the played audio information.

[1508] This series of processing steps enables store staff to provide customers with product information quickly and accurately in response to their needs.

[1509] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1510] The system of this invention aims to analyze conversations with customers and provide optimal product information. This system includes terminals, servers, an emotion engine, and multiple technical means that utilize voice and text data.

[1511] Specifically, the system operates as follows:

[1512] Terminal operation

[1513] The terminal captures audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and transmitted to the server in an encrypted state.

[1514] Server Processing

[1515] The server first converts the received audio data into text data using speech recognition technology. Then, it analyzes the text data using natural language processing technology to extract keywords that indicate customer needs.

[1516] How the emotion engine works

[1517] Furthermore, the voice data is analyzed by an emotion engine to recognize the customer's emotional state (e.g., satisfaction, dissatisfaction, excitement, questioning, etc.). Based on this emotional information, the emotion engine adjusts the importance and priority of keywords in product recommendations.

[1518] Selection of the optimal product

[1519] The server searches its internal database by combining extracted keywords and sentiment information, and selects the optimal smartphone model using a predetermined algorithm. For example, conditions such as "camera performance," "battery life," and "price range" may be extracted as keywords. If the sentiment engine recognizes that the customer is excited, it may make suggestions, including bold promotions.

[1520] Product Information

[1521] The information from the selected smartphones is converted back into audio data using speech synthesis technology on the server, encrypted, and sent to the terminal. The terminal plays this received audio data and presents the product information to the staff verbally.

[1522] The emotion engine adjusts its tone according to the user's emotions; if the user appears anxious, it adjusts to provide information in a calmer voice. This is expected to further enhance user satisfaction.

[1523] Mode switching

[1524] The terminal can switch modes with the press of a button according to the staff's request. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[1525] Specific example

[1526] The following are some specific use cases.

[1527] 1. The user (customer) enters the store and begins a conversation with the staff.

[1528] User: "I'm looking for a smartphone with a great camera and long battery life."

[1529] 2. The device captures this conversation as audio data and sends it to the server.

[1530] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[1531] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[1532] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[1533] 6. The server converts the selection results into audio data and sends it to the terminal. At this time, the emotion engine applies a tone that corresponds to the user's excitement level.

[1534] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1535] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[1536] In this way, staff can quickly and accurately suggest smartphones that match the customer's needs and feelings. This system makes it possible to provide high-quality customer service without relying on the staff's knowledge.

[1537] ---

[1538] The above describes specific embodiments for carrying out the present invention.

[1539] The following describes the processing flow.

[1540] Step 1:

[1541] The terminal begins voice input when a staff member starts a conversation with a customer. The conversation is captured in real time as audio data in digital format.

[1542] Step 2:

[1543] The device encrypts the acquired audio data and securely transmits it to the server.

[1544] Step 3:

[1545] The server decrypts the received encrypted audio data. Then, it uses speech recognition technology to convert the audio data into text data.

[1546] Step 4:

[1547] The server analyzes the converted text data and uses natural language processing techniques to extract important keywords. For example, keywords such as "camera performance," "battery life," and "price" may be extracted.

[1548] Step 5:

[1549] The server simultaneously passes the voice data to the emotion engine, which analyzes the user's emotional state. The emotion engine recognizes the customer's emotions, such as excitement, satisfaction, dissatisfaction, or questioning, based on their voice tone and speaking style.

[1550] Step 6:

[1551] The server searches its internal database based on the extracted keywords and sentiment information obtained from the sentiment engine, and lists the relevant smartphone models.

[1552] Step 7:

[1553] The server selects the optimal smartphone model from the listed models using an evaluation algorithm that combines emotional information. For example, if the customer appears anxious, a model emphasizing ease of use will be selected.

[1554] Step 8:

[1555] The server generates information about the selected smartphone model as text data. For example, it might include information such as, "Model X features a high-performance camera and a large-capacity battery."

[1556] Step 9:

[1557] The server converts the generated text data into speech data using speech synthesis technology. During this process, an emotion engine adjusts the tone of the speech. For example, if the user is excited, the speech will be generated in an energetic tone; if they are calm, it will be generated in a gentle tone.

[1558] Step 10:

[1559] The server encrypts the audio data and sends it to the terminal.

[1560] Step 11:

[1561] The device decrypts the received encrypted audio data and prepares it for playback.

[1562] Step 12:

[1563] The device transmits the played audio data to staff via an intercom. This allows staff to make real-time recommendations for the most suitable smartphone for each customer.

[1564] Step 13:

[1565] The terminal allows staff to easily switch between smartphone suggestion mode and normal intercom mode by operating a button on the intercom.

[1566] Specific example:

[1567] 1. The user (customer) enters the store and begins a conversation with the staff.

[1568] User: "I'm looking for a smartphone with a great camera and long battery life."

[1569] 2. The device captures this conversation as audio data and sends it to the server.

[1570] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[1571] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[1572] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[1573] 6. The server converts the selection results into audio data, applies a tone corresponding to the user's excitement level using the emotion engine, and sends it to the terminal.

[1574] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1575] 8. The terminal can be switched from smartphone suggestion mode to normal intercom mode by the staff operating a button on the intercom.

[1576] In this way, staff can quickly and accurately suggest smartphones that match the customer's needs and feelings. This system makes it possible to provide high-quality customer service without relying on the staff's knowledge.

[1577] (Example 2)

[1578] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1579] Traditional customer service systems rely heavily on staff knowledge for product recommendations and responses, making it difficult to provide high-quality service that responds promptly to customer needs and emotions. Furthermore, manually converting voice data to text and performing sentiment analysis is time-consuming and inaccurate. This can lead to decreased customer satisfaction and negatively impact store sales. Additionally, insufficient security measures for transmitting and receiving voice data create a risk of information leakage.

[1580] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1581] In this invention, the server includes means for converting voice data into text data using voice recognition technology, means for analyzing the text data and extracting keywords using natural language processing technology, and means for analyzing the voice data using an emotion analysis engine to recognize the customer's emotional state. This makes it possible to quickly respond to customer needs and emotions and provide high-quality, secure customer service.

[1582] A "terminal" is a device that captures conversations between staff and customers and transmits the audio data to a server.

[1583] A "server" is a central processing unit that receives and analyzes audio data and provides optimal product information.

[1584] "Voice data" refers to digital signals acquired by the terminal, including the content of conversations between staff and customers.

[1585] "Encryption" refers to the process of transforming and protecting audio data for security purposes.

[1586] "Decryption" is the process of restoring encrypted audio data to its original state.

[1587] "Speech recognition technology" is a technology that converts speech data into text data.

[1588] "Text data" refers to data that represents conversation content as written information, converted using speech recognition technology.

[1589] "Natural language processing technology" is a general term for technologies that analyze text data and extract keywords.

[1590] "Keywords" are important terms that indicate customer needs.

[1591] An "emotion engine" is a system that analyzes voice data to recognize the customer's emotional state.

[1592] "Emotional state" refers to the mental state a customer expresses during a conversation, such as "satisfaction," "dissatisfaction," "excitement," or "question."

[1593] A "database" is a data storage system used to store product information and other necessary information.

[1594] An "algorithm" refers to a set of calculation and processing steps used to select the optimal product based on extracted keywords and emotional information.

[1595] "Speech synthesis technology" is a technology that converts text data into speech data.

[1596] "Playback" is the process by which a device outputs received audio data as audio.

[1597] "Mode" refers to the operating state of the device, and includes, for example, smartphone suggestion mode and normal intercom mode.

[1598] The system of the present invention is built to provide optimal product information through interaction with users (customers). This system includes terminals, servers, an emotion engine, and multiple technical means that utilize voice and text data.

[1599] Specifically, the system operates using the following hardware and software.

[1600] Terminal operation

[1601] The terminal captures audio in real time as store staff converse with customers. The hardware used includes a high-quality microphone with noise-canceling capabilities. The captured audio data is converted to a digital format using digital signal processing technology and transmitted to the server with enhanced security using AES-256 encryption technology.

[1602] Server Processing

[1603] The server receives encrypted audio data sent from the terminal and first decrypts it. Then, it converts the audio data into text data using the Google Speech-to-Text API. This conversion process removes noise and maintains quality. Next, it analyzes the text data using a natural language processing library (e.g., spaCy or NLTK) to extract keywords that indicate customer needs. The extracted keywords specifically describe the product features that the customer is looking for.

[1604] How the emotion engine works

[1605] The server has a built-in emotion engine (e.g., IBM Watson Tone Analyzer) that analyzes the customer's emotional state based on voice data. The emotion engine analyzes the customer's voice tone, speed, pitch, etc., and determines emotional states such as "satisfied," "dissatisfied," "excited," and "questionable." This emotional information is used to adjust the importance and priority of keywords in the product proposal process described later.

[1606] Selection of the optimal product

[1607] The server searches its internal database based on extracted keywords and sentiment information, and selects the optimal product using a pre-configured algorithm (e.g., weighting calculations and evaluation functions). For example, if a customer prioritizes "camera performance" and "battery life," and the sentiment engine detects excitement, a product proposal, including a bold promotion, will be made.

[1608] Product Information

[1609] Information on selected products is converted back into audio data by the server using the Google Text-to-Speech API, and an emotion engine sets a tone that matches the customer's emotions. The re-encrypted audio data is sent to the terminal. The terminal decrypts the received audio data and plays it back through the staff member's earphones or speaker, providing the product information. Text information is also displayed on the terminal's screen.

[1610] Mode switching

[1611] The terminal features an interface that can be operated according to the staff's requests, and allows for easy switching between smartphone suggestion mode and normal intercom mode.

[1612] Specific example

[1613] The following are some specific use cases.

[1614] 1. The user (customer) enters the store and begins a conversation with the staff.

[1615] User: "I'm looking for a smartphone with a great camera and long battery life."

[1616] 2. The device captures this conversation as audio data and sends it to the server. This is done using the built-in microphone to capture high-quality audio data, which is then encrypted.

[1617] 3. The server converts the audio data into text data using the Google Speech-to-Text API and performs noise filtering. Next, it uses natural language processing technology (e.g., spaCy) to extract keywords such as "camera performance" and "long battery life."

[1618] 4. An emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's voice data and recognizes that the user is in an excited state.

[1619] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm (e.g., a weighting algorithm). For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[1620] 6. The server converts the selection results into audio data using the Google Text-to-Speech API and sends it to the device. At this time, the emotion engine applies a tone that corresponds to the user's level of excitement.

[1621] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1622] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[1623] Examples of prompts for generative AI models

[1624] The following are examples of prompts for a generative AI model.

[1625] User: "I'm looking for a smartphone with a great camera and long battery life."

[1626] This system allows staff to quickly and accurately recommend smartphones that match the customer's needs and feelings. It also enables high-quality customer service without relying on the staff's knowledge.

[1627] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1628] Program processing flow

[1629] Step 1: The device acquires the audio.

[1630] The terminal captures conversations between staff and users (customers) in real time using its built-in microphone. The audio data is converted into a digital signal and encrypted at high speed.

[1631] Input: Audio of a conversation between staff and customer

[1632] Processing: Digital signal conversion and encryption of audio (using AES-256 encryption technology)

[1633] Output: Encrypted digital audio data

[1634] Step 2: The device sends the audio data to the server.

[1635] The device sends encrypted voice data to the server. This communication is conducted using a secure protocol.

[1636] Input: Encrypted digital audio data

[1637] Processing: Data transmission (using a secure communication protocol)

[1638] Output: Data arrives on the server.

[1639] Step 3: The server decrypts the audio data.

[1640] The server receives encrypted audio data sent from the terminal and decrypts it using AES-256 decryption technology.

[1641] Input: Encrypted digital audio data

[1642] Processing: Decoding of audio data (using AES-256 decoding technology)

[1643] Output: Decoded digital audio data

[1644] Step 4: The server converts the audio to text.

[1645] The server uses the Google Speech-to-Text API to convert the audio data into text data.

[1646] Input: Decoded digital audio data

[1647] Processing: Speech recognition (using Google Speech-to-Text API)

[1648] Output: Text data

[1649] Step 5: The server extracts the keywords.

[1650] The server uses natural language processing tools such as spaCy and NLTK to extract important keywords from text data.

[1651] Input: Text data

[1652] Processing: Natural language processing (keyword extraction, TF-IDF calculation, etc.)

[1653] Output: Extracted keywords

[1654] Step 6: The emotion engine analyzes emotions.

[1655] The server uses an emotion engine to analyze voice data and recognize the user's emotional state.

[1656] Input: Decoded digital audio data

[1657] Processing: Sentiment analysis (analyzes voice tone, word choice, speed, etc.)

[1658] Output: Emotional information (satisfied, dissatisfied, excited, questioned, etc.)

[1659] Step 7: Select the optimal product for the server.

[1660] The server searches its internal database based on the extracted keywords and sentiment information to select the most suitable product.

[1661] Input: Extracted keywords, sentiment information

[1662] Processing: Database search and evaluation algorithms (weighting calculation, similarity calculation, etc.)

[1663] Output: Optimal product information

[1664] Step 8: The server generates the audio data.

[1665] The server uses the Google Text-to-Speech API to convert selected product information into audio data. The tone is then adjusted by an emotion engine.

[1666] Input: Optimal product information

[1667] Processing: Text-to-speech (using Google Text-to-Speech API)

[1668] Output: Audio data

[1669] Step 9: The server sends encrypted audio data to the device.

[1670] The server re-encrypts the generated audio data and sends it to the terminal.

[1671] Input: Audio data

[1672] Processing: Encryption of audio data (using AES-256 encryption technology) and transmission.

[1673] Output: Encrypted audio data

[1674] Step 10: The device plays the audio data.

[1675] The terminal decodes the received audio data and plays it back through earphones or speakers worn by the staff. It also displays text information on its screen.

[1676] Input: Encrypted audio data

[1677] Processing: Decoding and playback of audio data, displaying text on the screen.

[1678] Output: Played audio, displayed text information

[1679] Step 11: The device switches modes.

[1680] The terminal's mode can be switched by staff operation. This allows for easy switching between smartphone suggestion mode and normal intercom mode.

[1681] Input: Staff operation (button press)

[1682] Processing: Mode switching

[1683] Output: New mode setting

[1684] Example of a prompt

[1685] The following are examples of prompts for a generative AI model.

[1686] User: "I'm looking for a smartphone with a great camera and long battery life."

[1687] (Application Example 2)

[1688] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1689] Conventional customer service systems often rely on subjective judgments from staff when recommending products, failing to adequately reflect customers' specific needs and emotional states. Furthermore, providing timely and appropriate product information is difficult, resulting in a lack of effective means to improve customer satisfaction. This invention aims to solve these problems by providing a system that delivers optimal product information in real time, based on customer needs and emotions.

[1690] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's emotional state using an emotion engine and adjusting the tone of product suggestions based on the emotional state; means for generating the tone of product suggestions as prompt sentences using a generation AI model; and means for extracting keywords using natural language processing technology and selecting the optimal product using a recommendation system. This makes it possible to make product suggestions in an appropriate tone that takes into account the customer's emotional state, and is expected to improve customer satisfaction.

[1691] A "terminal" is a device that acquires conversations with a user as audio data, sends that data to a server, and plays back information from the server.

[1692] A "server" is a central processing unit that converts audio data into text data, analyzes that text data to extract keywords, analyzes the user's emotional state using an emotion engine, selects the optimal product, and transmits that information to the terminal.

[1693] An "emotion engine" is an algorithm or software that can analyze a user's voice data and recognize their emotional state.

[1694] "Keywords" are important words and phrases extracted from user conversations that are useful when selecting or proposing products.

[1695] "Natural language processing technology" is a technique for analyzing text data and extracting meaningful keywords from it.

[1696] A "recommendation system" is an algorithm or software that searches a database based on extracted keywords and selects the most suitable product.

[1697] A "generative AI model" is a machine learning-based artificial intelligence model used to generate the tone of product proposals.

[1698] A "prompt sentence" is an input sentence used to generate an AI model that includes the tone of a product proposal.

[1699] "Means for switching modes" refers to methods for changing the operating mode of a device depending on its intended use, such as switching between smartphone suggestion mode and normal intercom mode.

[1700] A system for specifically implementing the present invention is described below. This system acquires conversations between customers and staff in real time and provides optimal product information by analyzing the content of the conversation and the customer's emotional state.

[1701] Terminal operation

[1702] The terminal first captures the audio in real time as staff members converse with customers. The captured audio data is converted into a digital signal and sent to the server in an encrypted state. Standard protocols such as SSL / TLS are used for this encryption.

[1703] Server Processing

[1704] The server first converts the received audio data into text data using speech recognition technology. Google's speech recognition API is used here. Next, this text data is analyzed using natural language processing technology to extract keywords that indicate customer needs. The Python nltk library is used for natural language processing.

[1705] Furthermore, the server uses an emotion engine to analyze the voice data and recognize the customer's emotional state (e.g., satisfaction, dissatisfaction, excitement, questioning, etc.). The emotion engine uses the Hugging Face transformers library.

[1706] Selection of the optimal product

[1707] The server searches its internal database based on extracted keywords and sentiment information, and selects the optimal product using a predetermined algorithm. During this process, the database search and recommendation system work in conjunction. The recommendation system utilizes the scikit-learn library.

[1708] The selected product information is then adjusted by an emotion engine to match the customer's emotional tone. For example, if the customer is excited, the suggestions will be presented in an energetic tone. These suggestions are generated as prompts based on a generative AI model.

[1709] Product Information

[1710] The information on the selected products is converted back into audio data by the server using speech synthesis technology. The pyttsx3 library is used for this speech synthesis. Next, the audio data is encrypted again and sent to the terminal. The terminal plays the received audio data and presents the product information to the staff verbally.

[1711] Mode switching

[1712] The device has a function that allows staff to switch modes with the press of a button, according to their request. This makes it easy to switch between smartphone suggestion mode and normal intercom mode.

[1713] Specific example

[1714] The following are some specific use cases.

[1715] 1. The user (customer) enters the store and begins a conversation with the staff.

[1716] User: "I'm looking for a smartphone with a great camera and long battery life."

[1717] 2. The device captures this conversation as audio data and sends it to the server.

[1718] 3. The server converts the audio data into text data and uses natural language processing technology to extract keywords such as "camera performance" and "long battery life."

[1719] 4. The emotion engine analyzes the user's voice data and recognizes that the user is in an excited state.

[1720] 5. The server searches the database for multiple smartphone models based on keywords and sentiment information, and selects the optimal model using an evaluation algorithm. For example, "Model X" might be selected as a smartphone with a high-performance camera and a large-capacity battery.

[1721] 6. The server converts the selection results into audio data and sends it to the terminal. At this time, the emotion engine applies a tone that corresponds to the user's excitement level.

[1722] 7. The device plays the received audio data and informs the staff that "Model X features a high-performance camera and a large-capacity battery."

[1723] 8. The staff member switches from smartphone suggestion mode to normal intercom mode by operating a button on the intercom.

[1724] Example of a prompt

[1725] For example, if a customer says, "I'm looking for a smartphone with a great camera and long battery life," the following prompt will be generated.

[1726] A customer says, "I'm looking for a smartphone with excellent camera performance and long battery life." Please recommend the best product for them. Use sentiment analysis to ensure your recommendation is delivered in an appropriate tone.

[1727] This system allows staff to quickly and accurately suggest products that match the customer's needs and feelings. This is expected to improve customer satisfaction.

[1728] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1729] Step 1:

[1730] The terminal acquires the conversation with the user as audio data in real time. The input is the audio data of the conversation between the user and the staff, and the output is the audio data converted into a digital signal.

[1731] Step 2:

[1732] The terminal encrypts the acquired audio data and sends it to the server. The input is audio data converted into a digital signal, and the output is encrypted audio data. Standard protocols such as SSL / TLS are used for encryption.

[1733] Step 3:

[1734] The server converts received audio data into text data using speech recognition technology. The input is encrypted audio data, and the output is text data. Google's speech recognition API is used for speech recognition.

[1735] Step 4:

[1736] The server analyzes text data using natural language processing techniques and extracts keywords. The input is text data obtained through speech recognition technology, and the output is the analyzed keywords. The NLTK library is used for natural language processing.

[1737] Step 5:

[1738] The server analyzes voice data using an emotion engine to recognize the user's emotional state. The input is voice data acquired in real time, and the output is the recognized emotional state. The Hugging Face transformers library is used as the emotion engine.

[1739] Step 6:

[1740] The server combines extracted keywords and sentiment information to search an internal database and selects the optimal product using an evaluation algorithm. The input is keywords and sentiment information, and the output is information on the selected product. The recommendation system uses the scikit-learn library.

[1741] Step 7:

[1742] The server uses a generative AI model to adjust the tone of product suggestions and generate prompt sentences. The input is selected product information and sentiment information, and the output is the prompt sentence. A pre-trained language model is used for the generative AI model.

[1743] Step 8:

[1744] The server converts product information back into audio data, encrypts it, and sends it to the terminal. The input is a prompt message, and the output is encrypted audio data. The pyttsx3 library is used for speech synthesis.

[1745] Step 9:

[1746] The terminal decrypts and plays back the received audio data. The input is encrypted audio data, and the output is the decrypted and played audio information. This audio information is presented to staff as a product suggestion via the terminal.

[1747] Step 10:

[1748] Staff members can switch between smartphone suggestion mode and normal intercom mode by operating a button on the device. The input is the button operation signal, and the output is the mode switching status of the device.

[1749] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1750] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1751] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1752] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1753] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1754] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1755] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1756] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1757] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1758] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1759] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1760] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1761] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1762] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1763] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1764] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1765] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1766] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1767] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1768] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1769] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1770] The following is further disclosed regarding the embodiments described above.

[1771] (Claim 1)

[1772] A means by which the terminal acquires the conversation with the user as audio data,

[1773] The terminal provides means for transmitting the acquired audio data to the server,

[1774] The server provides means for converting the audio data into text data,

[1775] The server has means for analyzing the text data and extracting keywords,

[1776] The server provides means for selecting the optimal product based on the extracted keywords,

[1777] The server converts the information of the selected product into audio data and transmits it to the terminal,

[1778] The terminal comprises means for playing the received audio data,

[1779] The aforementioned terminal has means for switching modes,

[1780] A system that includes this.

[1781] (Claim 2)

[1782] The system according to claim 1, wherein the terminal encrypts the audio data and transmits it to the server, and the server decrypts the audio data.

[1783] (Claim 3)

[1784] The system according to claim 1, wherein the server extracts keywords using natural language processing technology and selects the optimal product using a recommendation system.

[1785] "Example 1"

[1786] (Claim 1)

[1787] A means by which the terminal acquires the conversation with the user as audio data,

[1788] The terminal provides means for transmitting the acquired audio data to the server,

[1789] The server provides means for converting the audio data into text data,

[1790] The server has means for analyzing the text data and extracting keywords,

[1791] The server provides means for selecting the optimal product based on the extracted keywords,

[1792] The server converts the information of the selected product into audio data and transmits it to the terminal,

[1793] The terminal comprises means for playing the received audio data,

[1794] The aforementioned terminal has means for switching modes,

[1795] The aforementioned terminal has means for encrypting voice data,

[1796] The server has means for decrypting audio data,

[1797] The server includes means for converting speech into text data using a speech recognition system,

[1798] The server includes means for extracting keywords from text data using natural language processing technology,

[1799] The server has means for searching an internal database and selecting the optimal product using an evaluation algorithm,

[1800] The server includes means for converting product information into voice data using speech synthesis technology,

[1801] A system that includes this.

[1802] (Claim 2)

[1803] The system according to claim 1, wherein the terminal encrypts the audio data and transmits it to the server, and the server decrypts the audio data.

[1804] (Claim 3)

[1805] The system according to claim 1, wherein the server extracts keywords using natural language processing technology and selects the optimal product using a recommendation system.

[1806] "Application Example 1"

[1807] (Claim 1)

[1808] A means by which the terminal acquires the conversation with the user as audio data,

[1809] The terminal provides means for transmitting the acquired audio data to the server,

[1810] The server provides means for converting the audio data into text data,

[1811] The server has means for analyzing the text data and extracting keywords,

[1812] The server provides means for selecting the optimal product based on the extracted keywords,

[1813] The server converts the information of the selected product into audio data and transmits it to the terminal,

[1814] The server uses a generation AI model when converting product information into audio data,

[1815] The terminal comprises means for playing the received audio data,

[1816] The aforementioned terminal has means for switching modes,

[1817] A system that includes this.

[1818] (Claim 2)

[1819] The system according to claim 1, wherein the terminal encrypts the audio data and transmits it to the server, and the server decrypts the audio data.

[1820] (Claim 3)

[1821] The system according to claim 1, wherein the server extracts keywords using natural language processing technology and selects the optimal product using a recommendation system.

[1822] "Example 2 of combining an emotion engine"

[1823] (Claim 1)

[1824] A means by which the terminal acquires the conversation with the user as audio data,

[1825] The terminal provides means for encrypting the acquired audio data and transmitting it to the server,

[1826] The server includes means for decoding the audio data and converting it into text data using speech recognition technology,

[1827] The server includes means for analyzing the text data using natural language processing technology and extracting keywords,

[1828] The emotion engine includes means for analyzing the aforementioned voice data to recognize the customer's emotional state,

[1829] The server has means for searching a database based on the extracted keywords and sentiment information and selecting the optimal product.

[1830] The server converts the information of the selected product into voice data using speech synthesis technology and transmits it to the terminal.

[1831] The terminal comprises means for playing back the received audio data,

[1832] The aforementioned terminal has means for switching modes,

[1833] A system that includes this.

[1834] (Claim 2)

[1835] The system according to claim 1, wherein the terminal encrypts the audio data and transmits it to the server, and the server decrypts the audio data.

[1836] (Claim 3)

[1837] The system according to claim 1, wherein the server extracts keywords using natural language processing technology and selects the optimal product using a recommendation system.

[1838] "Application example 2 when combining with an emotional engine"

[1839] (Claim 1)

[1840] A means by which the terminal acquires the conversation with the user as audio data,

[1841] The terminal provides means for transmitting the acquired audio data to the server,

[1842] The server provides means for converting the audio data into text data,

[1843] The server has means for analyzing the text data and extracting keywords,

[1844] The server provides means for selecting the optimal product based on the extracted keywords,

[1845] The server converts the information of the selected product into audio data and transmits it to the terminal,

[1846] The terminal provides means for playing the received audio data,

[1847] The aforementioned terminal has means for switching modes,

[1848] A server analyzes the user's emotional state using an emotion engine and adjusts the tone of product suggestions based on the said emotional state.

[1849] The server provides means for generating a product proposal tone as a prompt sentence using a generated AI model,

[1850] A system that includes this.

[1851] (Claim 2)

[1852] The system according to claim 1, wherein the terminal encrypts the audio data and transmits it to the server, and the server decrypts the audio data.

[1853] (Claim 3)

[1854] The system according to claim 1, wherein the server extracts keywords using natural language processing technology and selects the optimal product using a recommendation system. [Explanation of Symbols]

[1855] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means by which the terminal acquires the conversation with the user as audio data, The terminal provides means for transmitting the acquired audio data to the server, The server provides means for converting the audio data into text data, The server has means for analyzing the text data and extracting keywords, The server provides means for selecting the optimal product based on the extracted keywords, The server converts the information of the selected product into audio data and transmits it to the terminal, The terminal comprises means for playing the received audio data, The aforementioned terminal has means for switching modes, A system that includes this.

2. The system according to claim 1, wherein the terminal encrypts the audio data and transmits it to the server, and the server decrypts the audio data.

3. The system according to claim 1, wherein the server extracts keywords using natural language processing technology and selects the optimal product using a recommendation system.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A