System

A system that collects and analyzes voice and image data to efficiently manage conversations and personal information addresses the challenge of data inefficiencies, allowing users to quickly access and review interactions.

JP2026017359APending Publication Date: 2026-02-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118141
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

AI Technical Summary

Technical Problem

The challenge of efficiently recording and managing conversation content and personal information is significant for individuals with many interactions, leading to inefficiencies and strained interpersonal relationships.

Method used

A system that collects language and image data using voice and image recognition, analyzes the data to extract conversation content and identify individuals, and stores it in a database for efficient retrieval using keywords.

Benefits of technology

Enables users to quickly search and review past conversations and personal information, improving work efficiency and interpersonal relationships by providing accurate and accessible data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017359000001_ABST
    Figure 2026017359000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting language data; means for collecting image data; means for analyzing the collected language data to extract dialogue content; means for analyzing the collected image data to identify a person; means for transmitting the analyzed language data and image data to a server; means for storing the transmitted data in a database for each user; means for searching the stored data based on a keyword; and means for displaying a search result on a user's terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, people experience a wide variety of conversations and contacts on a daily basis, but it is difficult to accurately remember the content of those conversations and the people involved. This burden is particularly great for sales professionals who come into contact with many people at work, and individuals with a wide range of friendships. If this problem is left unaddressed, it could lead to a decline in work efficiency and friction in interpersonal relationships. Therefore, there is a need for a system that can efficiently record and manage conversation content and personal information. [Means for solving the problem]

[0005] To solve this problem, the present invention provides the following means. First, language data and image data are collected using a terminal equipped with voice recognition and image recognition functions. The collected language data is analyzed to extract the content of the conversation. The image data is analyzed to identify the person. The analysis results are sent to a server and stored in a database for each user. When a user searches for data based on keywords, the stored data is searched and the results are displayed on the user's terminal. This allows users to use a system that allows them to easily review past conversations and personal information. In addition, the collected data is automatically indexed based on keywords, providing a means to improve search efficiency. Furthermore, the system also includes a means for comparing image data with an existing person database and automatically updating new person information. This allows users to always have the latest person information.

[0006] "Language Data" refers to conversational or audio information collected using speech recognition functionality.

[0007] "Image data" refers to visual information, particularly images and photographs, collected using image recognition technology, including faces.

[0008] "Analysis" refers to the process of processing collected data to extract dialogue content and personal information.

[0009] "Dialogue content" refers to the content of the conversation and keywords extracted from the voice data.

[0010] "Person identification" refers to the process of identifying a person from image data and matching it with an existing database.

[0011] "Server" refers to a central processing unit for storing analysis results and for retrieving and managing data.

[0012] "User" refers to an individual who uses the system.

[0013] "Database" refers to an information storage system for storing and indexing collected language and image data.

[0014] "Keywords" refer to specific words or phrases that users enter when searching for information.

[0015] "Searching" refers to the process of locating information in a database based on specific keywords.

[0016] "Indexing" refers to the process of organizing information and associating it with specific keywords in order to search the data efficiently.

[0017] "Matching" refers to the process of comparing collected image data with an existing database of people to identify matching people.

[0018] "Terminal" refers to a device that has the ability to collect and analyze audio and images. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] The present invention is realized as a system that can collect, analyze, store, and search language and image data. This system consists of a terminal with speech recognition and image recognition capabilities, a server that stores the data, and an application that allows users to access the data through an interface.

[0041] Audio and image collection

[0042] Language data collection

[0043] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology.

[0044] Image data collection

[0045] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, which is also temporarily stored in its internal memory. The camera takes high-resolution images and is capable of identifying people's faces.

[0046] Data analysis

[0047] Language Data Analysis

[0048] Voice recognition software is used to analyze the voice data collected by the device. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are temporarily stored on the device.

[0049] Image data analysis

[0050] Image recognition software is used to analyze the image data collected by the device. This software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also temporarily stored within the device.

[0051] Data transmission and storage

[0052] Sending data

[0053] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[0054] Data storage

[0055] The server stores the analysis data it receives in a database for each user. The database is provided with metadata such as date and time, geographical information, and keywords. The database is also designed to automatically generate indexes to improve search efficiency.

[0056] Finding and Viewing Data

[0057] Search by keyword

[0058] Users enter keywords they want to search for through a dedicated application or web portal, which refer to specific conversation content or personal information.

[0059] Searching for Data

[0060] The server searches the user's database for relevant data based on the keywords entered, filtering the data based on metadata such as date, time, and geographic location.

[0061] Displaying search results

[0062] The server sends the search results to the user's device, which displays the received search results, allowing the user to check past conversations and personal information.

[0063] Specific examples

[0064] Example 1: When you want to review the conversation after a business meeting

[0065] The device collects and analyzes audio and video during business negotiations in real time. After the negotiation is over, when the user searches for the keyword "business negotiation" in the application, the server finds and displays the relevant analysis data. The user can easily review the details of the negotiation and the other party's information.

[0066] Example 2: You want to remember the name of someone you met at an event

[0067] During the event, the device collects and analyzes audio and images. After the event is over, when a user searches for "event," the server finds relevant data and displays photos, names, conversations, etc. This allows users to easily recall the names and conversations of people they met at the event.

[0068] The above is an embodiment of the system according to the present invention. This system records many conversations and contacts and provides an environment in which necessary information can be easily searched and viewed.

[0069] The processing flow will be explained below.

[0070] Step 1:

[0071] The device collects the audio

[0072] The device's built-in microphone captures surrounding sounds in real time, and the captured audio data is temporarily stored in digital format in the device's internal memory.

[0073] Step 2:

[0074] The device collects the images

[0075] The camera on the device captures images within its field of view in real time, and the captured image data is temporarily stored in the internal memory frame by frame.

[0076] Step 3:

[0077] The device analyzes the audio data

[0078] The device's internal voice recognition software analyzes the temporarily stored voice data to extract the conversation content and keywords. The extracted information is converted into text format, and the analysis results are temporarily stored on the device.

[0079] Step 4:

[0080] The device analyzes the image data

[0081] The device's internal image recognition software analyzes the temporarily stored image data to detect faces. The detected faces are then compared with an existing database to identify specific people. The results of this analysis are also temporarily stored on the device.

[0082] Step 5:

[0083] The device sends the data to the server

[0084] The results of the audio and image analysis are packaged and sent to a server in an encrypted format, along with metadata such as date, time, and location information.

[0085] Step 6:

[0086] The server receives the data

[0087] The server receives the data sent from the device and stores it in a database for each user, where it is indexed for efficient searching.

[0088] Step 7:

[0089] The user enters a keyword

[0090] Users enter keywords for the content they want to search for through a dedicated application or web portal. Keywords refer to specific conversation content or personal information.

[0091] Step 8:

[0092] The server searches for the data

[0093] The server searches the user's database for data related to the entered keywords, and a search algorithm takes into account date, time, location, and indexing to filter out relevant data.

[0094] Step 9:

[0095] The server sends the search results

[0096] The filtered search results are sent to the user's device, and are provided in text, image, and audio formats.

[0097] Step 10:

[0098] Your device will display the search results.

[0099] The search results received by the user's device are displayed on the screen, allowing the user to visually check related conversation content and other party information.

[0100] The above are the specific processing steps of the system.

[0101] Example 1

[0102] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0103] Conventional data collection and analysis systems are inefficient in collecting, analyzing, storing, and searching language and image data, making it difficult for users to quickly obtain the information they need. Furthermore, the lack of quality improvement measures, such as noise filtering for voice data and high-resolution image capture, reduces the reliability of data and reduces search efficiency.

[0104] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0105] In this invention, the server includes means for collecting language data, means for collecting image data, means for analyzing the collected language data to extract important information and keywords, means for analyzing the collected image data to identify people, means for packaging the analyzed language data and image data, encrypting the data, and transmitting the packaged data to the server, means for storing the transmitted data in a database for each user and adding metadata, means for searching indexed data based on keywords, means for displaying search results on the user's device, means for applying noise filtering technology to clarify voice data, means for capturing image data at high resolution and performing facial recognition, means for a user to input search keywords through an application or web portal, and for the server to efficiently search for related data based on the index, and means for visually displaying search results, thereby improving the quality of collected data and enabling users to quickly search for and confirm the information they need.

[0106] "Means for collecting language data" refers to devices or software for capturing and storing speech as digital data.

[0107] "Means for collecting image data" refers to a device or software for capturing and storing visual information as a digital image.

[0108] "Means of analyzing collected language data to extract key information and keywords" refers to software or algorithms that convert speech data into text data and identify specific words or phrases within it.

[0109] "Means for analyzing collected image data to identify persons" refers to software or algorithms that recognize human faces in images and identify specific persons by matching them with existing databases.

[0110] "Means for packaging analyzed language data and image data, encrypting them, and transmitting them to a server" refers to a device or software for packaging the analyzed information into a single data package, encrypting it, and transmitting it to a server via a communications network.

[0111] "Means for storing the transmitted data in a database for each user and adding metadata" refers to a device or software for classifying data based on the identification information of each user, adding supplementary information such as date and time and geographical information, and storing the data in a database.

[0112] "Means for searching indexed data based on keywords" refers to algorithms or software for efficiently searching data within a database using specific keywords.

[0113] "Means for displaying search results on a user's terminal" refers to a device or software for visually displaying the search result data sent from the server on a user interface.

[0114] "Means for applying noise filtering techniques to make the audio data clear" refers to algorithms or software that remove background noise from the audio data to improve the sound quality.

[0115] "Means for capturing high-resolution image data and performing facial recognition" refers to a device or software that captures high-resolution photographs and identifies human faces from those images.

[0116] "Means for a user to input search keywords through an application or web portal and for the server to efficiently search for related data based on an index" refers to a device or software that allows a user to input keywords through an interface and for the server to quickly search for data related to those keywords.

[0117] "Means for visually displaying search results" refers to a device or software that visually displays the search results provided by the server on a screen so that the user can confirm the results.

[0118] This invention is a system that collects, analyzes, stores, and searches voice and image data. This system consists of a terminal equipped with voice recognition and image recognition functions, a server for storing data, and an application for user access.

[0119] Audio data collection

[0120] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. For example, this microphone is used to obtain high-quality recordings of audio during meetings. The collected audio data is temporarily stored in the device's internal memory. Noise filtering technology (e.g., noise cancellation software) is used to improve the quality of the audio data.

[0121] Image data collection

[0122] The device is equipped with a high-resolution camera that can capture images of what is within its field of view in real time. For example, this camera can be used to capture images of meeting scenes and participants' faces. The collected image data is temporarily stored in internal memory. Furthermore, facial recognition software can be used to identify people's faces from the captured images.

[0123] Data analysis

[0124] Dedicated analysis software is used to analyze the voice and image data collected by the device. For voice data, voice recognition software (e.g., Google Speech-to-Text API) is used to convert the data into text and extract important dialogue and keywords. For image data, image recognition software (e.g., OpenCV) is used to detect people's faces and match them with existing information in a database. The analysis results are temporarily stored inside the device.

[0125] Data transmission and storage

[0126] The device packages the analyzed language data and image data, encrypts them, and sends them to the server. The data is encrypted using technology such as AES encryption. The server stores the received data in a database for each user, adding metadata such as date and time, geographic information, and keywords. It also automatically generates an index to improve search efficiency.

[0127] Searching for Data

[0128] Users can search for data by entering specific keywords through a dedicated application or web portal. The server then references the index based on the entered keywords and efficiently searches for related data. For example, users can enter keywords such as "business negotiation" or "event" to find related analysis data.

[0129] Displaying search results

[0130] The server sends the search results to the user's device, which then visually displays the results, allowing the user to check past interactions and personal information.

[0131] Specific examples

[0132] Example 1: When you want to review the conversation after a business meeting

[0133] The device collects and analyzes audio and video during the negotiation in real time. After the negotiation is over, when the user searches for the keyword "negotiation" in the application, the server finds the relevant analysis data and displays the results. This allows the user to easily review the details of the negotiation and the other party's information.

[0134] Example 2: You want to remember the name of someone you met at an event

[0135] During the event, the device collects and analyzes audio and images. After the event ends, when the user searches for "event" in the application, the server finds relevant data and displays photos, names, conversation details, etc. This allows the user to easily remember the names and conversation details of people they met at the event.

[0136] Prompt Sentence Examples

[0137] "I want to review the content of yesterday's business meeting."

[0138] "I forgot the name of the person I met at the event last week."

[0139] This concludes the description of the embodiment of the invention. This system improves the quality of data and enables users to quickly search and check the information they need.

[0140] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0141] Step 1: Collecting language and image data

[0142] The device activates a high-performance microphone and a high-resolution camera to capture the surrounding audio and video in real time. For example, it collects audio and video of participants in a meeting. This is the input, and the data is temporarily stored in the internal memory. The output is raw audio and video data.

[0143] Specific behavior:

[0144] The device activates the microphone and captures audio data.

[0145] The device activates the camera and captures image data.

[0146] The collected audio and image data is temporarily stored in the internal memory.

[0147] Step 2: Filtering the data

[0148] The collected data contains noise and unnecessary information, so the device applies noise filtering technology. The audio data is filtered to make it clear, and unnecessary parts of the image data are removed. This improves the quality of the data. The input is raw audio data and image data, and the output is clear audio data and high-quality image data.

[0149] Specific behavior:

[0150] The device uses noise cancelling software to remove noise from the audio data.

[0151] The terminal filters out unnecessary parts of the image data.

[0152] Step 3: Analyze the data

[0153] The device runs analysis software to convert the filtered voice data into text and extract key keywords and dialogue. Similarly, facial recognition software is used on image data to identify people. The input is clear voice data and high-quality image data, and the output is analyzed text data and facial recognition results.

[0154] Specific behavior:

[0155] The device converts the voice data into text using the Google Speech-to-Text API.

[0156] The device extracts important keywords from the text.

[0157] The device uses OpenCV to identify people from image data.

[0158] Step 4: Sending data

[0159] The device packages the analyzed language data and image data, encrypts them using AES encryption, and then sends them to the server. The input is the analyzed text data and face recognition results, and the output is an encrypted data packet.

[0160] Specific behavior:

[0161] The device packages the analytical data together.

[0162] The device encrypts the data using AES encryption.

[0163] The device sends the encrypted data to the server.

[0164] Step 5: Save your data

[0165] The server decrypts the received encrypted data and stores it in a database for each user. Metadata such as date, time, geographical information, and keywords are added to the analyzed data. An index is also generated to enable efficient data searches. The input is the encrypted data packet, and the output is the analyzed data stored in the database.

[0166] Specific behavior:

[0167] The server decrypts the encrypted data.

[0168] The server adds metadata to the data.

[0169] The server stores the analysis data in a database.

[0170] The server generates the index.

[0171] Step 6: Search for data

[0172] A user enters search keywords through a dedicated application or web portal. The server references the index and efficiently searches for relevant data. The input is the user's search keywords, and the output is the search result data.

[0173] Specific behavior:

[0174] The user enters search keywords in the application.

[0175] The server references the index to find the relevant data.

[0176] Step 7: Viewing search results

[0177] The server sends the search results to the user's device, which then visually displays the results. The user can review the search results and refer to past interactions and personal information. The input is the search result data, and the output is the visually displayed search results.

[0178] Specific behavior:

[0179] The server sends the search results to the user's terminal.

[0180] The device converts the search results into a display format and displays them to the user in an easy-to-view format.

[0181] (Application example 1)

[0182] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0183] Traditional customer service in brick-and-mortar stores lacks an efficient method for managing customer information when customers return or for referencing past conversations. This results in low accuracy in customer recognition and response, making it difficult to improve customer experience. It also makes it difficult to offer product suggestions or services based on past conversation information, resulting in missed opportunities. Therefore, it is necessary to establish a method for streamlining customer service and providing personalized services.

[0184] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0185] In this invention, the server includes means for collecting language data, means for collecting image data, means for analyzing the collected language data to extract the content of the conversation, means for analyzing the collected image data to identify individuals, means for transmitting the analyzed language data and image data to a remote server, means for storing the transmitted data in a database for each user, means for searching the stored data based on keywords, means for displaying the search results on the user's terminal, means for searching and displaying previous interactions when the user visits again, and means for suggesting product information and services based on the customer's conversation history. This makes it possible to improve the efficiency of customer service and provide personalized services.

[0186] "Language Data" means speech information collected using speech recognition technology.

[0187] "Image data" refers to visual information captured using a camera or other imaging device.

[0188] "Dialogue content" refers to the specific content of the conversation obtained by analyzing the collected language data.

[0189] An "individual" is a specific person who can be identified by analyzing image data.

[0190] A "remote server" is a data storage and processing server accessible over a network.

[0191] A "user-specific database" is a database for storing and managing unique data for each user.

[0192] "Keywords" are important words or phrases that represent specific information and are used to conduct a search.

[0193] "User Terminal" means the electronic device used by a User to access the System.

[0194] "Previous interactions" refers to records of past conversations and responses with customers.

[0195] "Product information" refers to information about the details and specifications of the products being sold.

[0196] "Services" refers to various conveniences and support provided to users.

[0197] The present invention is realized as a system that can collect, analyze, store, and search language and image data. This system is composed of the following elements:

[0198] Hardware and Software

[0199] 1. Device:

[0200] Smartphones (e.g., general smartphones, iPhones, etc.)

[0201] Smart glasses (e.g., Google Glass)

[0202] The device is equipped with a high-performance microphone and camera, allowing it to capture audio and images in real time.

[0203] 2. Software:

[0204] Voice recognition software (e.g., Google Cloud Speech-to-Text)

[0205] Image recognition software (e.g., OpenCV)

[0206] Database software (e.g. MongoDB)

[0207] Cloud infrastructure (e.g., Amazon Web Services, AWS, etc.)

[0208] Data collection and analysis

[0209] 1. Collecting language data

[0210] The device's microphone captures surrounding sounds in real time and temporarily stores the audio data in internal memory. Noise filtering technology is used to clear the audio data and make it available for analysis.

[0211] 2. Image data collection

[0212] The device's camera captures images within its field of view in real time and temporarily stores the image data in its internal memory. Based on the high-resolution camera images, it is possible to identify an individual's face.

[0213] 3. Data Analysis

[0214] The collected voice data is converted into text using voice recognition software, and the content of the conversation and keywords are extracted. The image data is analyzed using image recognition software to identify individuals. The results of these analyses are temporarily stored on the device.

[0215] Data transmission and storage

[0216] 1. Data transmission

[0217] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[0218] 2. Data storage

[0219] The analysis data received by the server is stored in a database for each user, with metadata such as date and time, geographical information, and keywords added, and an index is automatically generated.

[0220] Finding and Viewing Data

[0221] 1. Search by keyword

[0222] Users enter keywords they want to search for through a dedicated application or web portal. Keywords refer to specific conversation content or personal information. The server searches for relevant data from the user's database based on the entered keywords, and sends and displays the search results on the user's device.

[0223] 2. Response when you return

[0224] When a customer returns, the device has the ability to search and display previous interactions, allowing store staff to respond quickly based on past customer information.

[0225] Customized product information suggestions

[0226] This includes a means to suggest product information and services that reflect the customer's interaction history and preferences based on the analysis results. For example, if a customer has previously talked about a specific product, other related products can be suggested.

[0227] Specific examples

[0228] Consider an example of use in a cafe. When a customer visits the cafe for the first time and orders a cappuccino, the order details are recorded and analyzed by the terminal and sent to the server. When the customer returns later, the staff member can quickly search for the customer's information and smoothly re-offer the same by asking, "Would you like your usual cappuccino?"

[0229] Prompt Sentence Examples

[0230] "Design a system for a customer support application that searches for a customer's order history based on their previous visits and provides them promptly."

[0231] As a result, the present invention makes it possible to improve the efficiency of customer service and provide customized services.

[0232] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0233] Step 1:

[0234] The device (a smartphone or smart glasses) captures the surrounding audio and video in real time.

[0235] Input: Audio data collected by a microphone, image data collected by a camera.

[0236] Operation: The device temporarily stores audio data in its internal memory and performs noise filtering to obtain clear audio data. Image data is also temporarily stored in its internal memory.

[0237] Output: Noise-filtered audio data and temporarily stored image data.

[0238] Step 2:

[0239] The voice data collected by the device is analyzed using voice recognition software (e.g., Google Cloud Speech-to-Text) and converted into text data.

[0240] Input: Noise-filtered audio data.

[0241] How it works: Speech recognition software converts voice data into text and extracts dialogue and keywords.

[0242] Output: Parsed text data and extracted keywords.

[0243] Step 3:

[0244] The image data collected by the device is analyzed using image recognition software (e.g., OpenCV) to identify individuals.

[0245] Input: Temporarily saved image data.

[0246] How it works: Image recognition software analyzes image data, identifies individuals' faces, and checks them against an existing database of people to confirm a match.

[0247] Output: Result data that identifies the individual.

[0248] Step 4:

[0249] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[0250] Input: Parsed text data, extracted keywords, and result data of personal identification.

[0251] How it works: The device combines all this data into a single package, encrypts it, and sends it to the server.

[0252] Output: Encrypted data package.

[0253] Step 5:

[0254] The server analyzes the received data package and stores the information in a database for each user.

[0255] Input: Encrypted data package.

[0256] How it works: The server decrypts the data package and stores the analyzed language data and image data in a database for each user. When saving, it adds metadata such as date, time, geographical information, and keywords.

[0257] Output: Parsed data stored in a per-user database.

[0258] Step 6:

[0259] Users enter search keywords through a dedicated application or web portal, and the server searches for relevant data.

[0260] Input: The search keyword entered by the user.

[0261] How it works: The server searches for relevant data based on the search keywords in the user's database and filters it using metadata.

[0262] Output: The relevant data found.

[0263] Step 7:

[0264] The server sends the search results to the user's terminal, which displays the results.

[0265] Input: The relevant data found.

[0266] Operation: The server sends the search results to the user's device, which receives and displays the data.

[0267] Output: Search results displayed on the user's device.

[0268] Step 8:

[0269] When a customer visits the store again, the device will recognize the visit and search for and display the previous interaction.

[0270] Input: Personal information of returning customers.

[0271] How it works: The server searches a database of past interactions and sends them to the user's device, which then displays the information.

[0272] Output: Returning customer information and previous interactions displayed on the terminal.

[0273] Step 9:

[0274] Recommend appropriate product information and services based on the customer's interaction history.

[0275] Input: Customer interaction history.

[0276] How it works: The server analyzes past interaction history, compiles relevant product and service information, and sends it to the user's device. The device then displays the information and makes appropriate suggestions.

[0277] Output: Product information and service suggestions displayed on the device.

[0278] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0279] This invention is realized as a system that combines a system that can collect, analyze, store, and search language and image data with an emotion engine that recognizes user emotions. This system consists of a terminal equipped with speech recognition, image recognition, and emotion recognition functions, a server that stores data, and an application that allows users to access the data through an interface.

[0280] Audio and image collection

[0281] Language data collection

[0282] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology.

[0283] Image data collection

[0284] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, which is also temporarily stored in its internal memory. The camera takes high-resolution images and is capable of identifying people's faces.

[0285] emotion recognition

[0286] Emotion Recognition from Speech Data

[0287] The device sends the collected voice data to an emotion recognition engine, which analyzes the user's voice tone, pitch, rhythm, etc. to recognize emotions. The analysis results are saved along with the content of the conversation.

[0288] Emotion recognition from image data

[0289] The device sends the collected image data to an emotion recognition engine, which then analyzes the user's facial expressions and movements to recognize their emotions. The analysis results are then saved along with the person's identification information.

[0290] Data analysis

[0291] Language Data Analysis

[0292] Voice recognition software is used to analyze the voice data collected by the device. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are converted into text format and stored along with emotion data.

[0293] Image data analysis

[0294] Image recognition software is used to analyze the image data collected by the device. The software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also stored along with the emotion data.

[0295] Data transmission and storage

[0296] Sending data

[0297] The terminal packages the analyzed language data, image data, and emotion data and transmits them to the server in an encrypted format.

[0298] Data storage

[0299] The server stores the analyzed data it receives in a database for each user. The database is provided with metadata such as date and time, geographical information, emotional data, and keywords. The database is also designed to automatically generate indexes to improve search efficiency.

[0300] Finding and Viewing Data

[0301] Search by keyword

[0302] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[0303] Searching for Data

[0304] The server searches the user's database for relevant data based on the entered keywords, filtering the data based on date, time, geographical information, emotional data, and indexes.

[0305] Displaying search results

[0306] The server sends the search results to the user's device, which then displays the received search results, allowing the user to visually check the relevant conversation content, the other party's information, and their emotional state.

[0307] Specific examples

[0308] Example 1: Reviewing the conversation and emotions after a business meeting

[0309] The device collects and analyzes voice, image data, and emotional data during a business meeting in real time. After the meeting is over, when the user searches for the keyword "business meeting" in the application, the server finds and displays the relevant analysis data. The user can easily review the details of the meeting, information about the other party, and their own and the other party's emotional states.

[0310] Example 2: When you want to remember the name or feelings of someone you met at an event

[0311] During the event, the device collects and analyzes audio, image data, and emotional data. After the event is over, when the user searches for "event," the server finds relevant data and displays photos, names, conversations, and emotional states. This allows the user to easily recall the names, conversations, and emotional states of people they met at the event.

[0312] The above is an embodiment of a system based on the present invention that includes an emotion engine. This system records many conversations and interactions, manages emotional states, and provides an environment in which necessary information can be easily searched and viewed.

[0313] The processing flow will be explained below.

[0314] Step 1:

[0315] The device collects the audio

[0316] The device's built-in microphone captures surrounding sounds in real time, and the captured audio data is temporarily stored in digital format in the device's internal memory.

[0317] Step 2:

[0318] The device collects the images

[0319] The camera on the device captures images within its field of view in real time, and the captured image data is temporarily stored in the internal memory frame by frame.

[0320] Step 3:

[0321] The device analyzes the audio data

[0322] The device's internal voice recognition software analyzes the temporarily stored voice data to extract the conversation content and keywords. The extracted information is converted into text format, and the analysis results are temporarily stored on the device.

[0323] Step 4:

[0324] The device analyzes the image data

[0325] The device's internal image recognition software analyzes the temporarily stored image data to detect faces. The detected faces are then compared with an existing database to identify specific people. The results of this analysis are also temporarily stored on the device.

[0326] Step 5:

[0327] The device recognizes emotions from voice data

[0328] The device's emotion engine analyzes the voice data and recognizes the user's emotions by analyzing their tone, pitch, rhythm, etc. The emotion analysis results are also temporarily stored on the device.

[0329] Step 6:

[0330] The device recognizes emotions from image data

[0331] The device's emotion engine analyzes the image data and recognizes the user's emotions by analyzing their facial expressions and movements. The results of this emotion analysis are also temporarily stored on the device.

[0332] Step 7:

[0333] The device sends the data to the server

[0334] The results of the voice and image analysis and emotional data are packaged and sent to a server in an encrypted format, along with metadata such as date, time, and location information.

[0335] Step 8:

[0336] The server receives the data

[0337] The server receives the data sent from the device and stores it in a database for each user, where it is indexed for efficient searching.

[0338] Step 9:

[0339] The user enters a keyword

[0340] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[0341] Step 10:

[0342] The server searches for the data

[0343] The server searches the user's database to find data related to the entered keywords, and a search algorithm takes into account date, time, location, index, and sentiment data to filter out relevant data.

[0344] Step 11:

[0345] The server sends the search results

[0346] The filtered search results are sent to the user's device, and are provided in text, image, and audio formats.

[0347] Step 12:

[0348] Your device will display the search results.

[0349] The search results received by the user's device are displayed on the screen, allowing the user to visually check the relevant conversation content, information about the other party, and their emotional state.

[0350] The above are the specific processing steps of the system including the emotion engine.

[0351] Example 2

[0352] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0353] Conventional data collection and analysis systems lack emotion recognition capabilities, making it difficult to accurately grasp the emotional state of users and targets. Furthermore, when storing and searching large amounts of data, the lack of appropriate indexes and metadata makes it difficult to efficiently search and browse information. As a result, users are unable to easily review the content of conversations and related information, potentially leading to delays in productivity and decision-making.

[0354] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting language data, means for collecting image data, means for clearing the collected language data using noise filtering technology, means for capturing the collected image data at high resolution to identify a person, means for transmitting the language data to an emotion recognition engine to analyze voice tone, pitch, rhythm, etc., means for transmitting the image data to the emotion recognition engine to analyze facial expressions, slight movements, etc., means for analyzing the collected language data to extract dialogue content, means for analyzing the collected image data to identify a person, means for transmitting the analyzed language data and image data to the server, means for storing the transmitted data in a database for each user, means for generating an index by adding metadata such as date and time, geographic information, and emotion data to the stored data, means for searching the stored data based on keywords, and means for displaying search results on the user's terminal. This enables effective collection, analysis, and search of dialogue content and related information, including the user's emotional state.

[0355] "Language data" is data that includes the content of communication in audio or text format.

[0356] "Image data" refers to data including still images and moving images collected using a camera or other imaging device.

[0357] "Noise filtering technology" is a technology that removes background noise and unnecessary sounds from collected audio data.

[0358] "High resolution" refers to a resolution that allows images and videos to be displayed clearly down to the fine details.

[0359] An "emotion recognition engine" refers to software or hardware functionality that analyzes audio and image data to identify a user's emotional state.

[0360] "Dialogue content" refers to the content and meaning of the conversation extracted from audio data or text data.

[0361] "Identifying a person" means identifying the face or features of a specific individual from image data.

[0362] "Collected Data" refers to all data collected by the system, including language data and image data.

[0363] A "database" refers to a collection of data organized and stored for a specific purpose.

[0364] "Metadata" refers to data that contains information about the stored data, such as date and time, geographical information, and emotional data.

[0365] An "index" refers to a list or catalog created within a database to make data search and management more efficient.

[0366] A "keyword" refers to a specific word or string of characters that a user enters for searching or filtering.

[0367] This invention is a system that not only collects language data and image data, but also analyzes, stores, and searches that data, but also combines it with an emotion engine that recognizes user emotions. This system is composed of a terminal equipped with speech recognition, image recognition, and emotion recognition functions, a server that stores data, and an application that allows users to access the data through an interface. Specific embodiments are described below.

[0368] System Configuration

[0369] Audio and image collection

[0370] 1. Collecting language data

[0371] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology (e.g., Adobe Audition).

[0372] 2. Image data collection

[0373] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, takes high-resolution photos, and temporarily stores them in its internal memory. The camera has the ability to identify faces, which can be used to identify targets.

[0374] emotion recognition

[0375] 3. Emotion Recognition from Speech Data

[0376] The device sends the collected voice data to an emotion recognition engine (e.g., Affectiva SDK), which analyzes the voice tone, pitch, rhythm, etc. to recognize emotions. The analysis results are saved along with the content of the conversation.

[0377] 4. Emotion Recognition from Image Data

[0378] The image data collected by the device is sent to an emotion recognition engine, which analyzes the user's facial expressions and movements to recognize their emotions. The analysis results are saved along with the person's identification information.

[0379] Data analysis

[0380] 5. Language Data Analysis

[0381] The device uses voice recognition software (e.g., Google Speech-to-Text) to analyze the voice data collected. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are converted into text format and stored along with emotion data.

[0382] 6. Analysis of image data

[0383] Image recognition software (e.g., Amazon Rekognition) is used to analyze the image data collected by the device. This software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also stored along with the emotion data.

[0384] Data transmission and storage

[0385] 7. Data transmission

[0386] The terminal transmits the analyzed language data, image data, and emotion data to the server in an encrypted format.

[0387] 8. Data Retention

[0388] The server stores the analyzed data it receives in a database for each user. At this time, the database is given metadata such as date and time, geographical information, emotion data, and keywords, and an index is automatically generated based on this information.

[0389] Finding and Viewing Data

[0390] 9. Search by Keyword

[0391] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[0392] 10. Searching for Data

[0393] The server searches the user's database for relevant data based on the entered keywords, filtering the data based on date, time, geographical information, emotional data, and indexes.

[0394] 11. Display of search results

[0395] The server sends the search results to the user's device, which then displays the received search results, allowing the user to visually confirm the relevant conversation content, the other party's information, and their emotional state.

[0396] Specific examples

[0397] Example 1: Reviewing the conversation and emotions after a business meeting

[0398] The device collects and analyzes voice, image data, and emotional data during negotiations in real time. After the negotiation is over, when the user searches for the keyword "negotiation" in the application, the server finds and displays the relevant analysis data. The user can easily review details of the negotiation, information about the other party, and even their own and the other party's emotional states.

[0399] Example 2: When you want to remember the name or feelings of someone you met at an event

[0400] During the event, the device collects and analyzes audio, image data, and emotional data. After the event is over, when the user searches for "event," the server finds relevant data and displays photos, names, conversations, and emotional states. This allows the user to easily recall the names, conversations, and emotional states of people they met at the event.

[0401] An example of a prompt to actually enter

[0402] 1. "I want to review the conversation and emotional state of this business meeting."

[0403] 2. "Tell me how you would use this system to remember information about people you met at events."

[0404] This allows the system to effectively collect, analyze, store, search, and display dialogue content and related information for users, improving the user experience.

[0405] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0406] Step 1:

[0407] Language data collection

[0408] The device uses a high-performance microphone to capture surrounding audio. The input is ambient audio data, which the device collects in real time and temporarily stores in its internal memory. Then, noise filtering technology is used to remove background noise and unwanted sounds from the collected audio data. The output is clear audio data.

[0409] Step 2:

[0410] Image data collection

[0411] The device uses a camera to capture video within its field of view in real time. The input is video data of the surroundings, which the device temporarily stores in its internal memory. The image data is captured at high resolution and processed using a function to identify human faces. The output is the image data after face identification.

[0412] Step 3:

[0413] Emotion Recognition from Speech Data

[0414] The device sends the collected voice data to the emotion recognition engine. The input is clear voice data, which the emotion recognition engine analyzes and recognizes the user's emotional state by analyzing the voice tone, pitch, and rhythm. The output is the emotion analysis result.

[0415] Step 4:

[0416] Emotion recognition from image data

[0417] The device sends the collected image data to the emotion recognition engine. The input is image data after face identification, which the emotion recognition engine analyzes to recognize the user's emotional state by analyzing facial expressions and micro-movements. The output is the image emotion analysis results.

[0418] Step 5:

[0419] Language Data Analysis

[0420] The device uses voice recognition software to analyze the collected voice data. The input is clear voice data, which the voice recognition software converts into text. The software then extracts important dialogue and keywords from the text data. The output is the analyzed text data and keywords.

[0421] Step 6:

[0422] Image data analysis

[0423] Image recognition software is used to analyze the image data collected by the device. The input is image data after face identification, which the image recognition software analyzes to detect human faces and identify people by comparing them with information in an existing database. The output is the person recognition result.

[0424] Step 7:

[0425] Sending data

[0426] The terminal packages the analyzed language data, image data, and emotion data and transmits them in encrypted form to the server. The input is the analyzed language data, image data, and emotion data, which the terminal encrypts and transmits, and the output is the encrypted data transmitted to the server.

[0427] Step 8:

[0428] Data storage

[0429] The server stores the received data in a database for each user. The input is encrypted data sent to the server, which decrypts it and stores it in the database. Here, metadata such as date and time, geographic information, emotion data, and keywords are added, and an index is generated. The output is an indexed database entry.

[0430] Step 9:

[0431] Search by keyword

[0432] A user enters search keywords through a dedicated application or web portal. The input is the user-specified search keywords, which are sent to the server. The output is a search query based on the keywords.

[0433] Step 10:

[0434] Searching for Data

[0435] The server searches for relevant data from the user's database based on keywords. The input is a search query, and the server filters the relevant data based on this using date, time, geographic information, sentiment data, and indexes. The output is a dataset of search results.

[0436] Step 11:

[0437] Displaying search results

[0438] The server sends the search results to the user's device. The input is a dataset of search results, which is sent to the user's device. The user's device displays the search results based on the received data, allowing the user to visually confirm related conversation content, information about the other party, and their emotional state. The output is the displayed search results.

[0439] (Application example 2)

[0440] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0441] In traditional brick-and-mortar stores, it was difficult for store associates to understand customer interactions and their emotional state in real time, making it difficult to optimize the quality of service. Furthermore, there was a lack of effective means to make optimal suggestions and follow-ups based on customer emotions, which could result in lower customer satisfaction. Furthermore, traditional systems lacked the ability to provide real-time feedback on collected data, making it difficult for store associates to respond immediately.

[0442] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting language data, means for collecting image data, means for extracting dialogue content from the collected language data, means for identifying people from the collected image data, means for recognizing emotions from the collected language data and image data, means for transmitting the analyzed language data, image data, and emotion data to the server, means for storing the transmitted data in a database for each user, means for providing feedback on the stored data in real time, means for searching the stored data based on keywords, and means for displaying search results on the user's terminal. This allows store clerks to grasp the dialogue and emotional state of customers in real time and respond appropriately, improving the quality of service and increasing customer satisfaction.

[0443] The "means for collecting language data" refers to a device or method that uses a microphone or the like to capture surrounding sounds and collect them as language data.

[0444] "Means for collecting image data" refers to a device or method that uses a camera or the like to capture images within the field of view and collects them as image data.

[0445] The "means for extracting dialogue content from collected language data" refers to a device or method that converts collected language data into text using speech recognition technology and extracts important dialogue content and keywords.

[0446] "Means for identifying a person from collected image data" refers to a device or method that uses image recognition technology to identify a person from collected image data and compare it with a database.

[0447] The "means for recognizing emotions from collected language data and image data" refers to a device or method for recognizing the emotional state of a user by analyzing voice tone, facial expressions, etc.

[0448] The "means for transmitting analyzed language data, image data, and emotion data to a server" refers to a device or method for packaging the analyzed data and transmitting it to a server via a network.

[0449] The "means for storing transmitted data in a database for each user" refers to a device or method for sorting and storing the analysis data received by the server for each user.

[0450] "Means for providing real-time feedback of stored data" refers to a device or method that immediately processes the analyzed and stored data and displays it in real-time through a user interface.

[0451] The term "means for searching stored data based on keywords" refers to a device or method for efficiently searching related information from a stored database based on keywords entered by a user.

[0452] "Means for displaying search results on the user's terminal" refers to a device or method for transmitting searched data to the user's terminal and visually displaying it through an interface.

[0453] This invention is a system that enables store clerks in brick-and-mortar stores to grasp the conversations and emotional state of customers in real time and optimize the quality of service. The system consists of smart glasses, a voice recognition engine, an image recognition engine, an emotion recognition engine, a server, cloud storage, and a real-time display. This system allows the details of conversations with customers and the emotional state of customers to be grasped instantly, enabling the provision of appropriate services.

[0454] Hardware and Software Configuration

[0455] Smart Glasses

[0456] The smart glasses are equipped with high-performance microphones and cameras that collect voice and image data from customer interactions in real time.

[0457] Speech Recognition Engine

[0458] The speech recognition engine analyzes the collected voice data and converts it into text. This analysis is performed using, for example, the Google Speech Recognition API.

[0459] Image Recognition Engine

[0460] The image recognition engine analyzes the collected image data and identifies the customer's face using tools such as OpenCV and DeepFace.

[0461] Emotion Recognition Engine

[0462] The emotion recognition engine analyzes collected voice and image data to recognize the customer's emotional state, using the Emotion API and machine learning models for emotion recognition.

[0463] Server and Cloud Storage

[0464] The collected data and analysis results are sent in encrypted form to a server and stored in a per-user database using cloud storage services such as AWS or Google Cloud Storage.

[0465] Real-time Display

[0466] Store associates receive real-time feedback through the smart glasses' display, allowing them to respond immediately to the customer's immediate needs and emotional state.

[0467] Processing flow explanation

[0468] 1. Audio and Image Collection:

[0469] The smart glasses collect customer interaction through a microphone and capture images of the surroundings through a camera, and these data are temporarily stored in the internal memory.

[0470] 2. Data Analysis:

[0471] The voice data is converted into text by a voice recognition engine, which then analyzes emotions from the tone and pitch of the voice, while the image data is analyzed by an image recognition engine to identify the customer's face and recognize their emotional state from their facial expressions.

[0472] 3. Data transmission and storage:

[0473] The analyzed data is encrypted and sent to a server, which then stores the data in cloud storage and organizes it for each user.

[0474] 4. Real-time feedback:

[0475] The smart glasses display allows store staff to see the customer's conversation and emotional state in real time, allowing them to provide optimal responses and suggestions based on this feedback.

[0476] Specific examples

[0477] 1. Know when a customer is interested in a specific product:

[0478] The program recognizes the customer's interests in real time and provides feedback to the store clerk, saying, "You are interested in this product."

[0479] 2. Recognize when a customer is unhappy:

[0480] The program detects a customer's dissatisfied facial expression and displays a warning to the store clerk saying, "The customer is dissatisfied."

[0481] Prompt Sentence Examples

[0482] An example of a prompt is: "When a customer asks a question about a particular product, the smart glasses analyze the conversation and the customer's facial expressions in real time and display feedback such as, 'It seems you are interested in this product. Further explanation is needed.'"

[0483] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0484] Step 1:

[0485] The device (smart glasses) collects the conversational voice with the customer using a microphone and captures the surrounding field of view using a camera. These data are temporarily stored in the internal memory. Voice data and image data are collected as input and saved in the internal memory as output. In concrete terms, the microphone captures the voice signal and the camera continuously takes images.

[0486] Step 2:

[0487] The device sends the collected voice data to a speech recognition engine, which converts it into text data. The input is voice data, which the speech recognition engine converts into text. Text data is generated as output. Specifically, the device uses the Google Speech Recognition API to convert the voice signal into text.

[0488] Step 3:

[0489] The device sends the collected image data to an image recognition engine to identify the customer's face. The input is image data, which the image recognition engine analyzes to identify the person. The output is information about the identified person. Specifically, the device uses OpenCV and DeepFace to analyze facial features and compare them with a database.

[0490] Step 4:

[0491] The device sends the collected voice and image data to the emotion recognition engine to analyze the customer's emotional state. The input is voice and image data, which the emotion recognition engine analyzes to obtain emotional data. Emotion data is generated as output. Specifically, the device uses the Emotion API to analyze voice tone and facial expressions.

[0492] Step 5:

[0493] The device encrypts the analyzed language data, image data, and emotion data and sends it to the server. The input is the analyzed data, which is encrypted and sent to the server via the network. The output is the data sent to the server. Specifically, the device sends the data using an encryption protocol such as SSL / TLS.

[0494] Step 6:

[0495] The server stores the analysis data it receives in a database for each user. The input is the analyzed data, which is organized for each user and stored in the database. The output is the analysis data stored in the database. Specifically, the server organizes and stores the data using AWS or Google Cloud Storage.

[0496] Step 7:

[0497] The server sends analysis data for feedback to the device in real time. The input is analysis data obtained from the user's database, which is sent to the device in real time. The output is feedback that is displayed immediately on the device. Specifically, the server sends data immediately using a protocol such as WebSocket.

[0498] Step 8:

[0499] Based on the feedback displayed on the terminal in real time, the user (store clerk) responds appropriately to the customer. The input is feedback data from the terminal, and the user provides services based on this. The output is the result of providing an appropriate service. In concrete terms, the user checks the display on the smart glasses and engages in a dialogue based on the feedback content.

[0500] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0501] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0502] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0503] [Second embodiment]

[0504] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0505] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0506] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0507] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0508] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0509] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0510] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0511] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0512] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0513] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0514] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0515] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0516] The present invention is realized as a system that can collect, analyze, store, and search language and image data. This system consists of a terminal with speech recognition and image recognition capabilities, a server that stores the data, and an application that allows users to access the data through an interface.

[0517] Audio and image collection

[0518] Language data collection

[0519] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology.

[0520] Image data collection

[0521] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, which is also temporarily stored in its internal memory. The camera takes high-resolution images and is capable of identifying people's faces.

[0522] Data analysis

[0523] Language Data Analysis

[0524] Voice recognition software is used to analyze the voice data collected by the device. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are temporarily stored on the device.

[0525] Image data analysis

[0526] Image recognition software is used to analyze the image data collected by the device. This software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also temporarily stored within the device.

[0527] Data transmission and storage

[0528] Sending data

[0529] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[0530] Data storage

[0531] The server stores the analysis data it receives in a database for each user. The database is provided with metadata such as date and time, geographical information, and keywords. The database is also designed to automatically generate indexes to improve search efficiency.

[0532] Finding and Viewing Data

[0533] Search by keyword

[0534] Users enter keywords they want to search for through a dedicated application or web portal, which refer to specific conversation content or personal information.

[0535] Searching for Data

[0536] The server searches the user's database for relevant data based on the keywords entered, filtering the data based on metadata such as date, time, and geographic location.

[0537] Displaying search results

[0538] The server sends the search results to the user's device, which displays the received search results, allowing the user to check past conversations and personal information.

[0539] Specific examples

[0540] Example 1: When you want to review the conversation after a business meeting

[0541] The device collects and analyzes audio and video during business negotiations in real time. After the negotiation is over, when the user searches for the keyword "business negotiation" in the application, the server finds and displays the relevant analysis data. The user can easily review the details of the negotiation and the other party's information.

[0542] Example 2: You want to remember the name of someone you met at an event

[0543] During the event, the device collects and analyzes audio and images. After the event is over, when a user searches for "event," the server finds relevant data and displays photos, names, conversations, etc. This allows users to easily recall the names and conversations of people they met at the event.

[0544] The above is an embodiment of the system according to the present invention. This system records many conversations and contacts and provides an environment in which necessary information can be easily searched and viewed.

[0545] The processing flow will be explained below.

[0546] Step 1:

[0547] The device collects the audio

[0548] The device's built-in microphone captures surrounding sounds in real time, and the captured audio data is temporarily stored in digital format in the device's internal memory.

[0549] Step 2:

[0550] The device collects the images

[0551] The camera on the device captures images within its field of view in real time, and the captured image data is temporarily stored in the internal memory frame by frame.

[0552] Step 3:

[0553] The device analyzes the audio data

[0554] The device's internal voice recognition software analyzes the temporarily stored voice data to extract the conversation content and keywords. The extracted information is converted into text format, and the analysis results are temporarily stored on the device.

[0555] Step 4:

[0556] The device analyzes the image data

[0557] The device's internal image recognition software analyzes the temporarily stored image data to detect faces. The detected faces are then compared with an existing database to identify specific people. The results of this analysis are also temporarily stored on the device.

[0558] Step 5:

[0559] The device sends the data to the server

[0560] The results of the audio and image analysis are packaged and sent to a server in an encrypted format, along with metadata such as date, time, and location information.

[0561] Step 6:

[0562] The server receives the data

[0563] The server receives the data sent from the device and stores it in a database for each user, where it is indexed for efficient searching.

[0564] Step 7:

[0565] The user enters a keyword

[0566] Users enter keywords for the content they want to search for through a dedicated application or web portal. Keywords refer to specific conversation content or personal information.

[0567] Step 8:

[0568] The server searches for the data

[0569] The server searches the user's database for data related to the entered keywords, and a search algorithm takes into account date, time, location, and indexing to filter out relevant data.

[0570] Step 9:

[0571] The server sends the search results

[0572] The filtered search results are sent to the user's device, and are provided in text, image, and audio formats.

[0573] Step 10:

[0574] Your device will display the search results.

[0575] The search results received by the user's device are displayed on the screen, allowing the user to visually check related conversation content and other party information.

[0576] The above are the specific processing steps of the system.

[0577] Example 1

[0578] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0579] Conventional data collection and analysis systems are inefficient in collecting, analyzing, storing, and searching language and image data, making it difficult for users to quickly obtain the information they need. Furthermore, the lack of quality improvement measures, such as noise filtering for voice data and high-resolution image capture, reduces the reliability of data and reduces search efficiency.

[0580] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0581] In this invention, the server includes means for collecting language data, means for collecting image data, means for analyzing the collected language data to extract important information and keywords, means for analyzing the collected image data to identify people, means for packaging the analyzed language data and image data, encrypting the data, and transmitting the packaged data to the server, means for storing the transmitted data in a database for each user and adding metadata, means for searching indexed data based on keywords, means for displaying search results on the user's device, means for applying noise filtering technology to clarify voice data, means for capturing image data at high resolution and performing facial recognition, means for a user to input search keywords through an application or web portal, and for the server to efficiently search for related data based on the index, and means for visually displaying search results, thereby improving the quality of collected data and enabling users to quickly search for and confirm the information they need.

[0582] "Means for collecting language data" refers to devices or software for capturing and storing speech as digital data.

[0583] "Means for collecting image data" refers to a device or software for capturing and storing visual information as a digital image.

[0584] "Means of analyzing collected language data to extract key information and keywords" refers to software or algorithms that convert speech data into text data and identify specific words or phrases within it.

[0585] "Means for analyzing collected image data to identify persons" refers to software or algorithms that recognize human faces in images and identify specific persons by matching them with existing databases.

[0586] "Means for packaging analyzed language data and image data, encrypting them, and transmitting them to a server" refers to a device or software for packaging the analyzed information into a single data package, encrypting it, and transmitting it to a server via a communications network.

[0587] "Means for storing the transmitted data in a database for each user and adding metadata" refers to a device or software for classifying data based on the identification information of each user, adding supplementary information such as date and time and geographical information, and storing the data in a database.

[0588] "Means for searching indexed data based on keywords" refers to algorithms or software for efficiently searching data within a database using specific keywords.

[0589] "Means for displaying search results on a user's terminal" refers to a device or software for visually displaying the search result data sent from the server on a user interface.

[0590] "Means for applying noise filtering techniques to make the audio data clear" refers to algorithms or software that remove background noise from the audio data to improve the sound quality.

[0591] "Means for capturing high-resolution image data and performing facial recognition" refers to a device or software that captures high-resolution photographs and identifies human faces from those images.

[0592] "Means for a user to input search keywords through an application or web portal and for the server to efficiently search for related data based on an index" refers to a device or software that allows a user to input keywords through an interface and for the server to quickly search for data related to those keywords.

[0593] "Means for visually displaying search results" refers to a device or software that visually displays the search results provided by the server on a screen so that the user can confirm the results.

[0594] This invention is a system that collects, analyzes, stores, and searches voice and image data. This system consists of a terminal equipped with voice recognition and image recognition functions, a server for storing data, and an application for user access.

[0595] Audio data collection

[0596] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. For example, this microphone is used to obtain high-quality recordings of audio during meetings. The collected audio data is temporarily stored in the device's internal memory. Noise filtering technology (e.g., noise cancellation software) is used to improve the quality of the audio data.

[0597] Image data collection

[0598] The device is equipped with a high-resolution camera that can capture images of what is within its field of view in real time. For example, this camera can be used to capture images of meeting scenes and participants' faces. The collected image data is temporarily stored in internal memory. Furthermore, facial recognition software can be used to identify people's faces from the captured images.

[0599] Data analysis

[0600] Dedicated analysis software is used to analyze the voice and image data collected by the device. For voice data, voice recognition software (e.g., Google Speech-to-Text API) is used to convert the data into text and extract important dialogue and keywords. For image data, image recognition software (e.g., OpenCV) is used to detect people's faces and match them with existing information in a database. The analysis results are temporarily stored inside the device.

[0601] Data transmission and storage

[0602] The device packages the analyzed language data and image data, encrypts them, and sends them to the server. The data is encrypted using technology such as AES encryption. The server stores the received data in a database for each user, adding metadata such as date and time, geographic information, and keywords. It also automatically generates an index to improve search efficiency.

[0603] Searching for Data

[0604] Users can search for data by entering specific keywords through a dedicated application or web portal. The server then references the index based on the entered keywords and efficiently searches for related data. For example, users can enter keywords such as "business negotiation" or "event" to find related analysis data.

[0605] Displaying search results

[0606] The server sends the search results to the user's device, which then visually displays the results, allowing the user to check past interactions and personal information.

[0607] Specific examples

[0608] Example 1: When you want to review the conversation after a business meeting

[0609] The device collects and analyzes audio and video during the negotiation in real time. After the negotiation is over, when the user searches for the keyword "negotiation" in the application, the server finds the relevant analysis data and displays the results. This allows the user to easily review the details of the negotiation and the other party's information.

[0610] Example 2: You want to remember the name of someone you met at an event

[0611] During the event, the device collects and analyzes audio and images. After the event ends, when the user searches for "event" in the application, the server finds relevant data and displays photos, names, conversation details, etc. This allows the user to easily remember the names and conversation details of people they met at the event.

[0612] Prompt Sentence Examples

[0613] "I want to review the content of yesterday's business meeting."

[0614] "I forgot the name of the person I met at the event last week."

[0615] This concludes the description of the embodiment of the invention. This system improves the quality of data and enables users to quickly search and check the information they need.

[0616] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0617] Step 1: Collecting language and image data

[0618] The device activates a high-performance microphone and a high-resolution camera to capture the surrounding audio and video in real time. For example, it collects audio and video of participants in a meeting. This is the input, and the data is temporarily stored in the internal memory. The output is raw audio and video data.

[0619] Specific behavior:

[0620] The device activates the microphone and captures audio data.

[0621] The device activates the camera and captures image data.

[0622] The collected audio and image data is temporarily stored in the internal memory.

[0623] Step 2: Filtering the data

[0624] The collected data contains noise and unnecessary information, so the device applies noise filtering technology. The audio data is filtered to make it clear, and unnecessary parts of the image data are removed. This improves the quality of the data. The input is raw audio data and image data, and the output is clear audio data and high-quality image data.

[0625] Specific behavior:

[0626] The device uses noise cancelling software to remove noise from the audio data.

[0627] The terminal filters out unnecessary parts of the image data.

[0628] Step 3: Analyze the data

[0629] The device runs analysis software to convert the filtered voice data into text and extract key keywords and dialogue. Similarly, facial recognition software is used on image data to identify people. The input is clear voice data and high-quality image data, and the output is analyzed text data and facial recognition results.

[0630] Specific behavior:

[0631] The device converts the voice data into text using the Google Speech-to-Text API.

[0632] The device extracts important keywords from the text.

[0633] The device uses OpenCV to identify people from image data.

[0634] Step 4: Sending data

[0635] The device packages the analyzed language data and image data, encrypts them using AES encryption, and then sends them to the server. The input is the analyzed text data and face recognition results, and the output is an encrypted data packet.

[0636] Specific behavior:

[0637] The device packages the analytical data together.

[0638] The device encrypts the data using AES encryption.

[0639] The device sends the encrypted data to the server.

[0640] Step 5: Save your data

[0641] The server decrypts the received encrypted data and stores it in a database for each user. Metadata such as date, time, geographical information, and keywords are added to the analyzed data. An index is also generated to enable efficient data searches. The input is the encrypted data packet, and the output is the analyzed data stored in the database.

[0642] Specific behavior:

[0643] The server decrypts the encrypted data.

[0644] The server adds metadata to the data.

[0645] The server stores the analysis data in a database.

[0646] The server generates the index.

[0647] Step 6: Search for data

[0648] A user enters search keywords through a dedicated application or web portal. The server references the index and efficiently searches for relevant data. The input is the user's search keywords, and the output is the search result data.

[0649] Specific behavior:

[0650] The user enters search keywords in the application.

[0651] The server references the index to find the relevant data.

[0652] Step 7: Viewing search results

[0653] The server sends the search results to the user's device, which then visually displays the results. The user can review the search results and refer to past interactions and personal information. The input is the search result data, and the output is the visually displayed search results.

[0654] Specific behavior:

[0655] The server sends the search results to the user's terminal.

[0656] The device converts the search results into a display format and displays them to the user in an easy-to-view format.

[0657] (Application example 1)

[0658] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0659] Traditional customer service in brick-and-mortar stores lacks an efficient method for managing customer information when customers return or for referencing past conversations. This results in low accuracy in customer recognition and response, making it difficult to improve customer experience. It also makes it difficult to offer product suggestions or services based on past conversation information, resulting in missed opportunities. Therefore, it is necessary to establish a method for streamlining customer service and providing personalized services.

[0660] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0661] In this invention, the server includes means for collecting language data, means for collecting image data, means for analyzing the collected language data to extract the content of the conversation, means for analyzing the collected image data to identify individuals, means for transmitting the analyzed language data and image data to a remote server, means for storing the transmitted data in a database for each user, means for searching the stored data based on keywords, means for displaying the search results on the user's terminal, means for searching and displaying previous interactions when the user visits again, and means for suggesting product information and services based on the customer's conversation history. This makes it possible to improve the efficiency of customer service and provide personalized services.

[0662] "Language Data" means speech information collected using speech recognition technology.

[0663] "Image data" refers to visual information captured using a camera or other imaging device.

[0664] "Dialogue content" refers to the specific content of the conversation obtained by analyzing the collected language data.

[0665] An "individual" is a specific person who can be identified by analyzing image data.

[0666] A "remote server" is a data storage and processing server accessible over a network.

[0667] A "user-specific database" is a database for storing and managing unique data for each user.

[0668] "Keywords" are important words or phrases that represent specific information and are used to conduct a search.

[0669] "User Terminal" means the electronic device used by a User to access the System.

[0670] "Previous interactions" refers to records of past conversations and responses with customers.

[0671] "Product information" refers to information about the details and specifications of the products being sold.

[0672] "Services" refers to various conveniences and support provided to users.

[0673] The present invention is realized as a system that can collect, analyze, store, and search language and image data. This system is composed of the following elements:

[0674] Hardware and Software

[0675] 1. Device:

[0676] Smartphones (e.g., general smartphones, iPhones, etc.)

[0677] Smart glasses (e.g., Google Glass)

[0678] The device is equipped with a high-performance microphone and camera, allowing it to capture audio and images in real time.

[0679] 2. Software:

[0680] Voice recognition software (e.g., Google Cloud Speech-to-Text)

[0681] Image recognition software (e.g., OpenCV)

[0682] Database software (e.g. MongoDB)

[0683] Cloud infrastructure (e.g., Amazon Web Services, AWS, etc.)

[0684] Data collection and analysis

[0685] 1. Collecting language data

[0686] The device's microphone captures surrounding sounds in real time and temporarily stores the audio data in internal memory. Noise filtering technology is used to clear the audio data and make it available for analysis.

[0687] 2. Image data collection

[0688] The device's camera captures images within its field of view in real time and temporarily stores the image data in its internal memory. Based on the high-resolution camera images, it is possible to identify an individual's face.

[0689] 3. Data Analysis

[0690] The collected voice data is converted into text using voice recognition software, and the content of the conversation and keywords are extracted. The image data is analyzed using image recognition software to identify individuals. The results of these analyses are temporarily stored on the device.

[0691] Data transmission and storage

[0692] 1. Data transmission

[0693] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[0694] 2. Data storage

[0695] The analysis data received by the server is stored in a database for each user, with metadata such as date and time, geographical information, and keywords added, and an index is automatically generated.

[0696] Finding and Viewing Data

[0697] 1. Search by keyword

[0698] Users enter keywords they want to search for through a dedicated application or web portal. Keywords refer to specific conversation content or personal information. The server searches for relevant data from the user's database based on the entered keywords, and sends and displays the search results on the user's device.

[0699] 2. Response when you return

[0700] When a customer returns, the device has the ability to search and display previous interactions, allowing store staff to respond quickly based on past customer information.

[0701] Customized product information suggestions

[0702] This includes a means to suggest product information and services that reflect the customer's interaction history and preferences based on the analysis results. For example, if a customer has previously talked about a specific product, other related products can be suggested.

[0703] Specific examples

[0704] Consider an example of use in a cafe. When a customer visits the cafe for the first time and orders a cappuccino, the order details are recorded and analyzed by the terminal and sent to the server. When the customer returns later, the staff member can quickly search for the customer's information and smoothly re-offer the same by asking, "Would you like your usual cappuccino?"

[0705] Prompt Sentence Examples

[0706] "Design a system for a customer support application that searches for a customer's order history based on their previous visits and provides them promptly."

[0707] As a result, the present invention makes it possible to improve the efficiency of customer service and provide customized services.

[0708] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0709] Step 1:

[0710] The device (a smartphone or smart glasses) captures the surrounding audio and video in real time.

[0711] Input: Audio data collected by a microphone, image data collected by a camera.

[0712] Operation: The device temporarily stores audio data in its internal memory and performs noise filtering to obtain clear audio data. Image data is also temporarily stored in its internal memory.

[0713] Output: Noise-filtered audio data and temporarily stored image data.

[0714] Step 2:

[0715] The voice data collected by the device is analyzed using voice recognition software (e.g., Google Cloud Speech-to-Text) and converted into text data.

[0716] Input: Noise-filtered audio data.

[0717] How it works: Speech recognition software converts voice data into text and extracts dialogue and keywords.

[0718] Output: Parsed text data and extracted keywords.

[0719] Step 3:

[0720] The image data collected by the device is analyzed using image recognition software (e.g., OpenCV) to identify individuals.

[0721] Input: Temporarily saved image data.

[0722] How it works: Image recognition software analyzes image data, identifies individuals' faces, and checks them against an existing database of people to confirm a match.

[0723] Output: Result data that identifies the individual.

[0724] Step 4:

[0725] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[0726] Input: Parsed text data, extracted keywords, and result data of personal identification.

[0727] How it works: The device combines all this data into a single package, encrypts it, and sends it to the server.

[0728] Output: Encrypted data package.

[0729] Step 5:

[0730] The server analyzes the received data package and stores the information in a database for each user.

[0731] Input: Encrypted data package.

[0732] How it works: The server decrypts the data package and stores the analyzed language data and image data in a database for each user. When saving, it adds metadata such as date, time, geographical information, and keywords.

[0733] Output: Parsed data stored in a per-user database.

[0734] Step 6:

[0735] Users enter search keywords through a dedicated application or web portal, and the server searches for relevant data.

[0736] Input: The search keyword entered by the user.

[0737] How it works: The server searches for relevant data based on the search keywords in the user's database and filters it using metadata.

[0738] Output: The relevant data found.

[0739] Step 7:

[0740] The server sends the search results to the user's terminal, which displays the results.

[0741] Input: The relevant data found.

[0742] Operation: The server sends the search results to the user's device, which receives and displays the data.

[0743] Output: Search results displayed on the user's device.

[0744] Step 8:

[0745] When a customer visits the store again, the device will recognize the visit and search for and display the previous interaction.

[0746] Input: Personal information of returning customers.

[0747] How it works: The server searches a database of past interactions and sends them to the user's device, which then displays the information.

[0748] Output: Returning customer information and previous interactions displayed on the terminal.

[0749] Step 9:

[0750] Recommend appropriate product information and services based on the customer's interaction history.

[0751] Input: Customer interaction history.

[0752] How it works: The server analyzes past interaction history, compiles relevant product and service information, and sends it to the user's device. The device then displays the information and makes appropriate suggestions.

[0753] Output: Product information and service suggestions displayed on the device.

[0754] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0755] This invention is realized as a system that combines a system that can collect, analyze, store, and search language and image data with an emotion engine that recognizes user emotions. This system consists of a terminal equipped with speech recognition, image recognition, and emotion recognition functions, a server that stores data, and an application that allows users to access the data through an interface.

[0756] Audio and image collection

[0757] Language data collection

[0758] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology.

[0759] Image data collection

[0760] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, which is also temporarily stored in its internal memory. The camera takes high-resolution images and is capable of identifying people's faces.

[0761] emotion recognition

[0762] Emotion Recognition from Speech Data

[0763] The device sends the collected voice data to an emotion recognition engine, which analyzes the user's voice tone, pitch, rhythm, etc. to recognize emotions. The analysis results are saved along with the content of the conversation.

[0764] Emotion recognition from image data

[0765] The device sends the collected image data to an emotion recognition engine, which then analyzes the user's facial expressions and movements to recognize their emotions. The analysis results are then saved along with the person's identification information.

[0766] Data analysis

[0767] Language Data Analysis

[0768] Voice recognition software is used to analyze the voice data collected by the device. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are converted into text format and stored along with emotion data.

[0769] Image data analysis

[0770] Image recognition software is used to analyze the image data collected by the device. The software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also stored along with the emotion data.

[0771] Data transmission and storage

[0772] Sending data

[0773] The terminal packages the analyzed language data, image data, and emotion data and transmits them to the server in an encrypted format.

[0774] Data storage

[0775] The server stores the analyzed data it receives in a database for each user. The database is provided with metadata such as date and time, geographical information, emotional data, and keywords. The database is also designed to automatically generate indexes to improve search efficiency.

[0776] Finding and Viewing Data

[0777] Search by keyword

[0778] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[0779] Searching for Data

[0780] The server searches the user's database for relevant data based on the entered keywords, filtering the data based on date, time, geographical information, emotional data, and indexes.

[0781] Displaying search results

[0782] The server sends the search results to the user's device, which then displays the received search results, allowing the user to visually check the relevant conversation content, the other party's information, and their emotional state.

[0783] Specific examples

[0784] Example 1: Reviewing the conversation and emotions after a business meeting

[0785] The device collects and analyzes voice, image data, and emotional data during a business meeting in real time. After the meeting is over, when the user searches for the keyword "business meeting" in the application, the server finds and displays the relevant analysis data. The user can easily review the details of the meeting, information about the other party, and their own and the other party's emotional states.

[0786] Example 2: When you want to remember the name or feelings of someone you met at an event

[0787] During the event, the device collects and analyzes audio, image data, and emotional data. After the event is over, when the user searches for "event," the server finds relevant data and displays photos, names, conversations, and emotional states. This allows the user to easily recall the names, conversations, and emotional states of people they met at the event.

[0788] The above is an embodiment of a system based on the present invention that includes an emotion engine. This system records many conversations and interactions, manages emotional states, and provides an environment in which necessary information can be easily searched and viewed.

[0789] The processing flow will be explained below.

[0790] Step 1:

[0791] The device collects the audio

[0792] The device's built-in microphone captures surrounding sounds in real time, and the captured audio data is temporarily stored in digital format in the device's internal memory.

[0793] Step 2:

[0794] The device collects the images

[0795] The camera on the device captures images within its field of view in real time, and the captured image data is temporarily stored in the internal memory frame by frame.

[0796] Step 3:

[0797] The device analyzes the audio data

[0798] The device's internal voice recognition software analyzes the temporarily stored voice data to extract the conversation content and keywords. The extracted information is converted into text format, and the analysis results are temporarily stored on the device.

[0799] Step 4:

[0800] The device analyzes the image data

[0801] The device's internal image recognition software analyzes the temporarily stored image data to detect faces. The detected faces are then compared with an existing database to identify specific people. The results of this analysis are also temporarily stored on the device.

[0802] Step 5:

[0803] The device recognizes emotions from voice data

[0804] The device's emotion engine analyzes the voice data and recognizes the user's emotions by analyzing their tone, pitch, rhythm, etc. The emotion analysis results are also temporarily stored on the device.

[0805] Step 6:

[0806] The device recognizes emotions from image data

[0807] The device's emotion engine analyzes the image data and recognizes the user's emotions by analyzing their facial expressions and movements. The results of this emotion analysis are also temporarily stored on the device.

[0808] Step 7:

[0809] The device sends the data to the server

[0810] The results of the voice and image analysis and emotional data are packaged and sent to a server in an encrypted format, along with metadata such as date, time, and location information.

[0811] Step 8:

[0812] The server receives the data

[0813] The server receives the data sent from the device and stores it in a database for each user, where it is indexed for efficient searching.

[0814] Step 9:

[0815] The user enters a keyword

[0816] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[0817] Step 10:

[0818] The server searches for the data

[0819] The server searches the user's database to find data related to the entered keywords, and a search algorithm takes into account date, time, location, index, and sentiment data to filter out relevant data.

[0820] Step 11:

[0821] The server sends the search results

[0822] The filtered search results are sent to the user's device, and are provided in text, image, and audio formats.

[0823] Step 12:

[0824] Your device will display the search results.

[0825] The search results received by the user's device are displayed on the screen, allowing the user to visually check the relevant conversation content, information about the other party, and their emotional state.

[0826] The above are the specific processing steps of the system including the emotion engine.

[0827] Example 2

[0828] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0829] Conventional data collection and analysis systems lack emotion recognition capabilities, making it difficult to accurately grasp the emotional state of users and targets. Furthermore, when storing and searching large amounts of data, the lack of appropriate indexes and metadata makes it difficult to efficiently search and browse information. As a result, users are unable to easily review the content of conversations and related information, potentially leading to delays in productivity and decision-making.

[0830] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting language data, means for collecting image data, means for clearing the collected language data using noise filtering technology, means for capturing the collected image data at high resolution to identify a person, means for transmitting the language data to an emotion recognition engine to analyze voice tone, pitch, rhythm, etc., means for transmitting the image data to the emotion recognition engine to analyze facial expressions, slight movements, etc., means for analyzing the collected language data to extract dialogue content, means for analyzing the collected image data to identify a person, means for transmitting the analyzed language data and image data to the server, means for storing the transmitted data in a database for each user, means for generating an index by adding metadata such as date and time, geographic information, and emotion data to the stored data, means for searching the stored data based on keywords, and means for displaying search results on the user's terminal. This enables effective collection, analysis, and search of dialogue content and related information, including the user's emotional state.

[0831] "Language data" is data that includes the content of communication in audio or text format.

[0832] "Image data" refers to data including still images and moving images collected using a camera or other imaging device.

[0833] "Noise filtering technology" is a technology that removes background noise and unnecessary sounds from collected audio data.

[0834] "High resolution" refers to a resolution that allows images and videos to be displayed clearly down to the fine details.

[0835] An "emotion recognition engine" refers to software or hardware functionality that analyzes audio and image data to identify a user's emotional state.

[0836] "Dialogue content" refers to the content and meaning of the conversation extracted from audio data or text data.

[0837] "Identifying a person" means identifying the face or features of a specific individual from image data.

[0838] "Collected Data" refers to all data collected by the system, including language data and image data.

[0839] A "database" refers to a collection of data organized and stored for a specific purpose.

[0840] "Metadata" refers to data that contains information about the stored data, such as date and time, geographical information, and emotional data.

[0841] An "index" refers to a list or catalog created within a database to make data search and management more efficient.

[0842] A "keyword" refers to a specific word or string of characters that a user enters for searching or filtering.

[0843] This invention is a system that not only collects language data and image data, but also analyzes, stores, and searches that data, but also combines it with an emotion engine that recognizes user emotions. This system is composed of a terminal equipped with speech recognition, image recognition, and emotion recognition functions, a server that stores data, and an application that allows users to access the data through an interface. Specific embodiments are described below.

[0844] System Configuration

[0845] Audio and image collection

[0846] 1. Collecting language data

[0847] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology (e.g., Adobe Audition).

[0848] 2. Image data collection

[0849] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, takes high-resolution photos, and temporarily stores them in its internal memory. The camera has the ability to identify faces, which can be used to identify targets.

[0850] emotion recognition

[0851] 3. Emotion Recognition from Speech Data

[0852] The device sends the collected voice data to an emotion recognition engine (e.g., Affectiva SDK), which analyzes the voice tone, pitch, rhythm, etc. to recognize emotions. The analysis results are saved along with the content of the conversation.

[0853] 4. Emotion Recognition from Image Data

[0854] The image data collected by the device is sent to an emotion recognition engine, which analyzes the user's facial expressions and movements to recognize their emotions. The analysis results are saved along with the person's identification information.

[0855] Data analysis

[0856] 5. Language Data Analysis

[0857] The device uses voice recognition software (e.g., Google Speech-to-Text) to analyze the voice data collected. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are converted into text format and stored along with emotion data.

[0858] 6. Analysis of image data

[0859] Image recognition software (e.g., Amazon Rekognition) is used to analyze the image data collected by the device. This software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also stored along with the emotion data.

[0860] Data transmission and storage

[0861] 7. Data transmission

[0862] The terminal transmits the analyzed language data, image data, and emotion data to the server in an encrypted format.

[0863] 8. Data Retention

[0864] The server stores the analyzed data it receives in a database for each user. At this time, the database is given metadata such as date and time, geographical information, emotion data, and keywords, and an index is automatically generated based on this information.

[0865] Finding and Viewing Data

[0866] 9. Search by Keyword

[0867] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[0868] 10. Searching for Data

[0869] The server searches the user's database for relevant data based on the entered keywords, filtering the data based on date, time, geographical information, emotional data, and indexes.

[0870] 11. Display of search results

[0871] The server sends the search results to the user's device, which then displays the received search results, allowing the user to visually confirm the relevant conversation content, the other party's information, and their emotional state.

[0872] Specific examples

[0873] Example 1: Reviewing the conversation and emotions after a business meeting

[0874] The device collects and analyzes voice, image data, and emotional data during negotiations in real time. After the negotiation is over, when the user searches for the keyword "negotiation" in the application, the server finds and displays the relevant analysis data. The user can easily review details of the negotiation, information about the other party, and even their own and the other party's emotional states.

[0875] Example 2: When you want to remember the name or feelings of someone you met at an event

[0876] During the event, the device collects and analyzes audio, image data, and emotional data. After the event is over, when the user searches for "event," the server finds relevant data and displays photos, names, conversations, and emotional states. This allows the user to easily recall the names, conversations, and emotional states of people they met at the event.

[0877] An example of a prompt to actually enter

[0878] 1. "I want to review the conversation and emotional state of this business meeting."

[0879] 2. "Tell me how you would use this system to remember information about people you met at events."

[0880] This allows the system to effectively collect, analyze, store, search, and display dialogue content and related information for users, improving the user experience.

[0881] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0882] Step 1:

[0883] Language data collection

[0884] The device uses a high-performance microphone to capture surrounding audio. The input is ambient audio data, which the device collects in real time and temporarily stores in its internal memory. Then, noise filtering technology is used to remove background noise and unwanted sounds from the collected audio data. The output is clear audio data.

[0885] Step 2:

[0886] Image data collection

[0887] The device uses a camera to capture video within its field of view in real time. The input is video data of the surroundings, which the device temporarily stores in its internal memory. The image data is captured at high resolution and processed using a function to identify human faces. The output is the image data after face identification.

[0888] Step 3:

[0889] Emotion Recognition from Speech Data

[0890] The device sends the collected voice data to the emotion recognition engine. The input is clear voice data, which the emotion recognition engine analyzes and recognizes the user's emotional state by analyzing the voice tone, pitch, and rhythm. The output is the emotion analysis result.

[0891] Step 4:

[0892] Emotion recognition from image data

[0893] The device sends the collected image data to the emotion recognition engine. The input is image data after face identification, which the emotion recognition engine analyzes to recognize the user's emotional state by analyzing facial expressions and micro-movements. The output is the image emotion analysis results.

[0894] Step 5:

[0895] Language Data Analysis

[0896] The device uses voice recognition software to analyze the collected voice data. The input is clear voice data, which the voice recognition software converts into text. The software then extracts important dialogue and keywords from the text data. The output is the analyzed text data and keywords.

[0897] Step 6:

[0898] Image data analysis

[0899] Image recognition software is used to analyze the image data collected by the device. The input is image data after face identification, which the image recognition software analyzes to detect human faces and identify people by comparing them with information in an existing database. The output is the person recognition result.

[0900] Step 7:

[0901] Sending data

[0902] The terminal packages the analyzed language data, image data, and emotion data and transmits them in encrypted form to the server. The input is the analyzed language data, image data, and emotion data, which the terminal encrypts and transmits, and the output is the encrypted data transmitted to the server.

[0903] Step 8:

[0904] Data storage

[0905] The server stores the received data in a database for each user. The input is encrypted data sent to the server, which decrypts it and stores it in the database. Here, metadata such as date and time, geographic information, emotion data, and keywords are added, and an index is generated. The output is an indexed database entry.

[0906] Step 9:

[0907] Search by keyword

[0908] A user enters search keywords through a dedicated application or web portal. The input is the user-specified search keywords, which are sent to the server. The output is a search query based on the keywords.

[0909] Step 10:

[0910] Searching for Data

[0911] The server searches for relevant data from the user's database based on keywords. The input is a search query, and the server filters the relevant data based on this using date, time, geographic information, sentiment data, and indexes. The output is a dataset of search results.

[0912] Step 11:

[0913] Displaying search results

[0914] The server sends the search results to the user's device. The input is a dataset of search results, which is sent to the user's device. The user's device displays the search results based on the received data, allowing the user to visually confirm related conversation content, information about the other party, and their emotional state. The output is the displayed search results.

[0915] (Application example 2)

[0916] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0917] In traditional brick-and-mortar stores, it was difficult for store associates to understand customer interactions and their emotional state in real time, making it difficult to optimize the quality of service. Furthermore, there was a lack of effective means to make optimal suggestions and follow-ups based on customer emotions, which could result in lower customer satisfaction. Furthermore, traditional systems lacked the ability to provide real-time feedback on collected data, making it difficult for store associates to respond immediately.

[0918] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting language data, means for collecting image data, means for extracting dialogue content from the collected language data, means for identifying people from the collected image data, means for recognizing emotions from the collected language data and image data, means for transmitting the analyzed language data, image data, and emotion data to the server, means for storing the transmitted data in a database for each user, means for providing feedback on the stored data in real time, means for searching the stored data based on keywords, and means for displaying search results on the user's terminal. This allows store clerks to grasp the dialogue and emotional state of customers in real time and respond appropriately, improving the quality of service and increasing customer satisfaction.

[0919] The "means for collecting language data" refers to a device or method that uses a microphone or the like to capture surrounding sounds and collect them as language data.

[0920] "Means for collecting image data" refers to a device or method that uses a camera or the like to capture images within the field of view and collects them as image data.

[0921] The "means for extracting dialogue content from collected language data" refers to a device or method that converts collected language data into text using speech recognition technology and extracts important dialogue content and keywords.

[0922] "Means for identifying a person from collected image data" refers to a device or method that uses image recognition technology to identify a person from collected image data and compare it with a database.

[0923] The "means for recognizing emotions from collected language data and image data" refers to a device or method for recognizing the emotional state of a user by analyzing voice tone, facial expressions, etc.

[0924] The "means for transmitting analyzed language data, image data, and emotion data to a server" refers to a device or method for packaging the analyzed data and transmitting it to a server via a network.

[0925] The "means for storing transmitted data in a database for each user" refers to a device or method for sorting and storing the analysis data received by the server for each user.

[0926] "Means for providing real-time feedback of stored data" refers to a device or method that immediately processes the analyzed and stored data and displays it in real-time through a user interface.

[0927] The term "means for searching stored data based on keywords" refers to a device or method for efficiently searching related information from a stored database based on keywords entered by a user.

[0928] "Means for displaying search results on the user's terminal" refers to a device or method for transmitting searched data to the user's terminal and visually displaying it through an interface.

[0929] This invention is a system that enables store clerks in brick-and-mortar stores to grasp the conversations and emotional state of customers in real time and optimize the quality of service. The system consists of smart glasses, a voice recognition engine, an image recognition engine, an emotion recognition engine, a server, cloud storage, and a real-time display. This system allows the details of conversations with customers and the emotional state of customers to be grasped instantly, enabling the provision of appropriate services.

[0930] Hardware and Software Configuration

[0931] Smart Glasses

[0932] The smart glasses are equipped with high-performance microphones and cameras that collect voice and image data from customer interactions in real time.

[0933] Speech Recognition Engine

[0934] The speech recognition engine analyzes the collected voice data and converts it into text. This analysis is performed using, for example, the Google Speech Recognition API.

[0935] Image Recognition Engine

[0936] The image recognition engine analyzes the collected image data and identifies the customer's face using tools such as OpenCV and DeepFace.

[0937] Emotion Recognition Engine

[0938] The emotion recognition engine analyzes collected voice and image data to recognize the customer's emotional state, using the Emotion API and machine learning models for emotion recognition.

[0939] Server and Cloud Storage

[0940] The collected data and analysis results are sent in encrypted form to a server and stored in a per-user database using cloud storage services such as AWS or Google Cloud Storage.

[0941] Real-time Display

[0942] Store associates receive real-time feedback through the smart glasses' display, allowing them to respond immediately to the customer's immediate needs and emotional state.

[0943] Processing flow explanation

[0944] 1. Audio and Image Collection:

[0945] The smart glasses collect customer interaction through a microphone and capture images of the surroundings through a camera, and these data are temporarily stored in the internal memory.

[0946] 2. Data Analysis:

[0947] The voice data is converted into text by a voice recognition engine, which then analyzes emotions from the tone and pitch of the voice, while the image data is analyzed by an image recognition engine to identify the customer's face and recognize their emotional state from their facial expressions.

[0948] 3. Data transmission and storage:

[0949] The analyzed data is encrypted and sent to a server, which then stores the data in cloud storage and organizes it for each user.

[0950] 4. Real-time feedback:

[0951] The smart glasses display allows store staff to see the customer's conversation and emotional state in real time, allowing them to provide optimal responses and suggestions based on this feedback.

[0952] Specific examples

[0953] 1. Know when a customer is interested in a specific product:

[0954] The program recognizes the customer's interests in real time and provides feedback to the store clerk, saying, "You are interested in this product."

[0955] 2. Recognize when a customer is unhappy:

[0956] The program detects a customer's dissatisfied facial expression and displays a warning to the store clerk saying, "The customer is dissatisfied."

[0957] Prompt Sentence Examples

[0958] An example of a prompt is: "When a customer asks a question about a particular product, the smart glasses analyze the conversation and the customer's facial expressions in real time and display feedback such as, 'It seems you are interested in this product. Further explanation is needed.'"

[0959] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0960] Step 1:

[0961] The device (smart glasses) collects the conversational voice with the customer using a microphone and captures the surrounding field of view using a camera. These data are temporarily stored in the internal memory. Voice data and image data are collected as input and saved in the internal memory as output. In concrete terms, the microphone captures the voice signal and the camera continuously takes images.

[0962] Step 2:

[0963] The device sends the collected voice data to a speech recognition engine, which converts it into text data. The input is voice data, which the speech recognition engine converts into text. Text data is generated as output. Specifically, the device uses the Google Speech Recognition API to convert the voice signal into text.

[0964] Step 3:

[0965] The device sends the collected image data to an image recognition engine to identify the customer's face. The input is image data, which the image recognition engine analyzes to identify the person. The output is information about the identified person. Specifically, the device uses OpenCV and DeepFace to analyze facial features and compare them with a database.

[0966] Step 4:

[0967] The device sends the collected voice and image data to the emotion recognition engine to analyze the customer's emotional state. The input is voice and image data, which the emotion recognition engine analyzes to obtain emotional data. Emotion data is generated as output. Specifically, the device uses the Emotion API to analyze voice tone and facial expressions.

[0968] Step 5:

[0969] The device encrypts the analyzed language data, image data, and emotion data and sends it to the server. The input is the analyzed data, which is encrypted and sent to the server via the network. The output is the data sent to the server. Specifically, the device sends the data using an encryption protocol such as SSL / TLS.

[0970] Step 6:

[0971] The server stores the analysis data it receives in a database for each user. The input is the analyzed data, which is organized for each user and stored in the database. The output is the analysis data stored in the database. Specifically, the server organizes and stores the data using AWS or Google Cloud Storage.

[0972] Step 7:

[0973] The server sends analysis data for feedback to the device in real time. The input is analysis data obtained from the user's database, which is sent to the device in real time. The output is feedback that is displayed immediately on the device. Specifically, the server sends data immediately using a protocol such as WebSocket.

[0974] Step 8:

[0975] Based on the feedback displayed on the terminal in real time, the user (store clerk) responds appropriately to the customer. The input is feedback data from the terminal, and the user provides services based on this. The output is the result of providing an appropriate service. In concrete terms, the user checks the display on the smart glasses and engages in a dialogue based on the feedback content.

[0976] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0977] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0978] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0979] [Third embodiment]

[0980] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0981] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0982] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0983] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0984] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0985] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0986] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0987] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0988] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0989] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0990] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0991] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0992] The present invention is realized as a system that can collect, analyze, store, and search language and image data. This system consists of a terminal with speech recognition and image recognition capabilities, a server that stores the data, and an application that allows users to access the data through an interface.

[0993] Audio and image collection

[0994] Language data collection

[0995] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology.

[0996] Image data collection

[0997] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, which is also temporarily stored in its internal memory. The camera takes high-resolution images and is capable of identifying people's faces.

[0998] Data analysis

[0999] Language Data Analysis

[1000] Voice recognition software is used to analyze the voice data collected by the device. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are temporarily stored on the device.

[1001] Image data analysis

[1002] Image recognition software is used to analyze the image data collected by the device. This software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also temporarily stored within the device.

[1003] Data transmission and storage

[1004] Sending data

[1005] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[1006] Data storage

[1007] The server stores the analysis data it receives in a database for each user. The database is provided with metadata such as date and time, geographical information, and keywords. The database is also designed to automatically generate indexes to improve search efficiency.

[1008] Finding and Viewing Data

[1009] Search by keyword

[1010] Users enter keywords they want to search for through a dedicated application or web portal, which refer to specific conversation content or personal information.

[1011] Searching for Data

[1012] The server searches the user's database for relevant data based on the keywords entered, filtering the data based on metadata such as date, time, and geographic location.

[1013] Displaying search results

[1014] The server sends the search results to the user's device, which displays the received search results, allowing the user to check past conversations and personal information.

[1015] Specific examples

[1016] Example 1: When you want to review the conversation after a business meeting

[1017] The device collects and analyzes audio and video during business negotiations in real time. After the negotiation is over, when the user searches for the keyword "business negotiation" in the application, the server finds and displays the relevant analysis data. The user can easily review the details of the negotiation and the other party's information.

[1018] Example 2: You want to remember the name of someone you met at an event

[1019] During the event, the device collects and analyzes audio and images. After the event is over, when a user searches for "event," the server finds relevant data and displays photos, names, conversations, etc. This allows users to easily recall the names and conversations of people they met at the event.

[1020] The above is an embodiment of the system according to the present invention. This system records many conversations and contacts and provides an environment in which necessary information can be easily searched and viewed.

[1021] The processing flow will be explained below.

[1022] Step 1:

[1023] The device collects the audio

[1024] The device's built-in microphone captures surrounding sounds in real time, and the captured audio data is temporarily stored in digital format in the device's internal memory.

[1025] Step 2:

[1026] The device collects the images

[1027] The camera on the device captures images within its field of view in real time, and the captured image data is temporarily stored in the internal memory frame by frame.

[1028] Step 3:

[1029] The device analyzes the audio data

[1030] The device's internal voice recognition software analyzes the temporarily stored voice data to extract the conversation content and keywords. The extracted information is converted into text format, and the analysis results are temporarily stored on the device.

[1031] Step 4:

[1032] The device analyzes the image data

[1033] The device's internal image recognition software analyzes the temporarily stored image data to detect faces. The detected faces are then compared with an existing database to identify specific people. The results of this analysis are also temporarily stored on the device.

[1034] Step 5:

[1035] The device sends the data to the server

[1036] The results of the audio and image analysis are packaged and sent to a server in an encrypted format, along with metadata such as date, time, and location information.

[1037] Step 6:

[1038] The server receives the data

[1039] The server receives the data sent from the device and stores it in a database for each user, where it is indexed for efficient searching.

[1040] Step 7:

[1041] The user enters a keyword

[1042] Users enter keywords for the content they want to search for through a dedicated application or web portal. Keywords refer to specific conversation content or personal information.

[1043] Step 8:

[1044] The server searches for the data

[1045] The server searches the user's database for data related to the entered keywords, and a search algorithm takes into account date, time, location, and indexing to filter out relevant data.

[1046] Step 9:

[1047] The server sends the search results

[1048] The filtered search results are sent to the user's device, and are provided in text, image, and audio formats.

[1049] Step 10:

[1050] Your device will display the search results.

[1051] The search results received by the user's device are displayed on the screen, allowing the user to visually check related conversation content and other party information.

[1052] The above are the specific processing steps of the system.

[1053] Example 1

[1054] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1055] Conventional data collection and analysis systems are inefficient in collecting, analyzing, storing, and searching language and image data, making it difficult for users to quickly obtain the information they need. Furthermore, the lack of quality improvement measures, such as noise filtering for voice data and high-resolution image capture, reduces the reliability of data and reduces search efficiency.

[1056] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1057] In this invention, the server includes means for collecting language data, means for collecting image data, means for analyzing the collected language data to extract important information and keywords, means for analyzing the collected image data to identify people, means for packaging the analyzed language data and image data, encrypting the data, and transmitting the packaged data to the server, means for storing the transmitted data in a database for each user and adding metadata, means for searching indexed data based on keywords, means for displaying search results on the user's device, means for applying noise filtering technology to clarify voice data, means for capturing image data at high resolution and performing facial recognition, means for a user to input search keywords through an application or web portal, and for the server to efficiently search for related data based on the index, and means for visually displaying search results, thereby improving the quality of collected data and enabling users to quickly search for and confirm the information they need.

[1058] "Means for collecting language data" refers to devices or software for capturing and storing speech as digital data.

[1059] "Means for collecting image data" refers to a device or software for capturing and storing visual information as a digital image.

[1060] "Means of analyzing collected language data to extract key information and keywords" refers to software or algorithms that convert speech data into text data and identify specific words or phrases within it.

[1061] "Means for analyzing collected image data to identify persons" refers to software or algorithms that recognize human faces in images and identify specific persons by matching them with existing databases.

[1062] "Means for packaging analyzed language data and image data, encrypting them, and transmitting them to a server" refers to a device or software for packaging the analyzed information into a single data package, encrypting it, and transmitting it to a server via a communications network.

[1063] "Means for storing the transmitted data in a database for each user and adding metadata" refers to a device or software for classifying data based on the identification information of each user, adding supplementary information such as date and time and geographical information, and storing the data in a database.

[1064] "Means for searching indexed data based on keywords" refers to algorithms or software for efficiently searching data within a database using specific keywords.

[1065] "Means for displaying search results on a user's terminal" refers to a device or software for visually displaying the search result data sent from the server on a user interface.

[1066] "Means for applying noise filtering techniques to make the audio data clear" refers to algorithms or software that remove background noise from the audio data to improve the sound quality.

[1067] "Means for capturing high-resolution image data and performing facial recognition" refers to a device or software that captures high-resolution photographs and identifies human faces from those images.

[1068] "Means for a user to input search keywords through an application or web portal and for the server to efficiently search for related data based on an index" refers to a device or software that allows a user to input keywords through an interface and for the server to quickly search for data related to those keywords.

[1069] "Means for visually displaying search results" refers to a device or software that visually displays the search results provided by the server on a screen so that the user can confirm the results.

[1070] This invention is a system that collects, analyzes, stores, and searches voice and image data. This system consists of a terminal equipped with voice recognition and image recognition functions, a server for storing data, and an application for user access.

[1071] Audio data collection

[1072] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. For example, this microphone is used to obtain high-quality recordings of audio during meetings. The collected audio data is temporarily stored in the device's internal memory. Noise filtering technology (e.g., noise cancellation software) is used to improve the quality of the audio data.

[1073] Image data collection

[1074] The device is equipped with a high-resolution camera that can capture images of what is within its field of view in real time. For example, this camera can be used to capture images of meeting scenes and participants' faces. The collected image data is temporarily stored in internal memory. Furthermore, facial recognition software can be used to identify people's faces from the captured images.

[1075] Data analysis

[1076] Dedicated analysis software is used to analyze the voice and image data collected by the device. For voice data, voice recognition software (e.g., Google Speech-to-Text API) is used to convert the data into text and extract important dialogue and keywords. For image data, image recognition software (e.g., OpenCV) is used to detect people's faces and match them with existing information in a database. The analysis results are temporarily stored inside the device.

[1077] Data transmission and storage

[1078] The device packages the analyzed language data and image data, encrypts them, and sends them to the server. The data is encrypted using technology such as AES encryption. The server stores the received data in a database for each user, adding metadata such as date and time, geographic information, and keywords. It also automatically generates an index to improve search efficiency.

[1079] Searching for Data

[1080] Users can search for data by entering specific keywords through a dedicated application or web portal. The server then references the index based on the entered keywords and efficiently searches for related data. For example, users can enter keywords such as "business negotiation" or "event" to find related analysis data.

[1081] Displaying search results

[1082] The server sends the search results to the user's device, which then visually displays the results, allowing the user to check past interactions and personal information.

[1083] Specific examples

[1084] Example 1: When you want to review the conversation after a business meeting

[1085] The device collects and analyzes audio and video during the negotiation in real time. After the negotiation is over, when the user searches for the keyword "negotiation" in the application, the server finds the relevant analysis data and displays the results. This allows the user to easily review the details of the negotiation and the other party's information.

[1086] Example 2: You want to remember the name of someone you met at an event

[1087] During the event, the device collects and analyzes audio and images. After the event ends, when the user searches for "event" in the application, the server finds relevant data and displays photos, names, conversation details, etc. This allows the user to easily remember the names and conversation details of people they met at the event.

[1088] Prompt Sentence Examples

[1089] "I want to review the content of yesterday's business meeting."

[1090] "I forgot the name of the person I met at the event last week."

[1091] This concludes the description of the embodiment of the invention. This system improves the quality of data and enables users to quickly search and check the information they need.

[1092] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1093] Step 1: Collecting language and image data

[1094] The device activates a high-performance microphone and a high-resolution camera to capture the surrounding audio and video in real time. For example, it collects audio and video of participants in a meeting. This is the input, and the data is temporarily stored in the internal memory. The output is raw audio and video data.

[1095] Specific behavior:

[1096] The device activates the microphone and captures audio data.

[1097] The device activates the camera and captures image data.

[1098] The collected audio and image data is temporarily stored in the internal memory.

[1099] Step 2: Filtering the data

[1100] The collected data contains noise and unnecessary information, so the device applies noise filtering technology. The audio data is filtered to make it clear, and unnecessary parts of the image data are removed. This improves the quality of the data. The input is raw audio data and image data, and the output is clear audio data and high-quality image data.

[1101] Specific behavior:

[1102] The device uses noise cancelling software to remove noise from the audio data.

[1103] The terminal filters out unnecessary parts of the image data.

[1104] Step 3: Analyze the data

[1105] The device runs analysis software to convert the filtered voice data into text and extract key keywords and dialogue. Similarly, facial recognition software is used on image data to identify people. The input is clear voice data and high-quality image data, and the output is analyzed text data and facial recognition results.

[1106] Specific behavior:

[1107] The device converts the voice data into text using the Google Speech-to-Text API.

[1108] The device extracts important keywords from the text.

[1109] The device uses OpenCV to identify people from image data.

[1110] Step 4: Sending data

[1111] The device packages the analyzed language data and image data, encrypts them using AES encryption, and then sends them to the server. The input is the analyzed text data and face recognition results, and the output is an encrypted data packet.

[1112] Specific behavior:

[1113] The device packages the analytical data together.

[1114] The device encrypts the data using AES encryption.

[1115] The device sends the encrypted data to the server.

[1116] Step 5: Save your data

[1117] The server decrypts the received encrypted data and stores it in a database for each user. Metadata such as date, time, geographical information, and keywords are added to the analyzed data. An index is also generated to enable efficient data searches. The input is the encrypted data packet, and the output is the analyzed data stored in the database.

[1118] Specific behavior:

[1119] The server decrypts the encrypted data.

[1120] The server adds metadata to the data.

[1121] The server stores the analysis data in a database.

[1122] The server generates the index.

[1123] Step 6: Search for data

[1124] A user enters search keywords through a dedicated application or web portal. The server references the index and efficiently searches for relevant data. The input is the user's search keywords, and the output is the search result data.

[1125] Specific behavior:

[1126] The user enters search keywords in the application.

[1127] The server references the index to find the relevant data.

[1128] Step 7: Viewing search results

[1129] The server sends the search results to the user's device, which then visually displays the results. The user can review the search results and refer to past interactions and personal information. The input is the search result data, and the output is the visually displayed search results.

[1130] Specific behavior:

[1131] The server sends the search results to the user's terminal.

[1132] The device converts the search results into a display format and displays them to the user in an easy-to-view format.

[1133] (Application example 1)

[1134] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1135] Traditional customer service in brick-and-mortar stores lacks an efficient method for managing customer information when customers return or for referencing past conversations. This results in low accuracy in customer recognition and response, making it difficult to improve customer experience. It also makes it difficult to offer product suggestions or services based on past conversation information, resulting in missed opportunities. Therefore, it is necessary to establish a method for streamlining customer service and providing personalized services.

[1136] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1137] In this invention, the server includes means for collecting language data, means for collecting image data, means for analyzing the collected language data to extract the content of the conversation, means for analyzing the collected image data to identify individuals, means for transmitting the analyzed language data and image data to a remote server, means for storing the transmitted data in a database for each user, means for searching the stored data based on keywords, means for displaying the search results on the user's terminal, means for searching and displaying previous interactions when the user visits again, and means for suggesting product information and services based on the customer's conversation history. This makes it possible to improve the efficiency of customer service and provide personalized services.

[1138] "Language Data" means speech information collected using speech recognition technology.

[1139] "Image data" refers to visual information captured using a camera or other imaging device.

[1140] "Dialogue content" refers to the specific content of the conversation obtained by analyzing the collected language data.

[1141] An "individual" is a specific person who can be identified by analyzing image data.

[1142] A "remote server" is a data storage and processing server accessible over a network.

[1143] A "user-specific database" is a database for storing and managing unique data for each user.

[1144] "Keywords" are important words or phrases that represent specific information and are used to conduct a search.

[1145] "User Terminal" means the electronic device used by a User to access the System.

[1146] "Previous interactions" refers to records of past conversations and responses with customers.

[1147] "Product information" refers to information about the details and specifications of the products being sold.

[1148] "Services" refers to various conveniences and support provided to users.

[1149] The present invention is realized as a system that can collect, analyze, store, and search language and image data. This system is composed of the following elements:

[1150] Hardware and Software

[1151] 1. Device:

[1152] Smartphones (e.g., general smartphones, iPhones, etc.)

[1153] Smart glasses (e.g., Google Glass)

[1154] The device is equipped with a high-performance microphone and camera, allowing it to capture audio and images in real time.

[1155] 2. Software:

[1156] Voice recognition software (e.g., Google Cloud Speech-to-Text)

[1157] Image recognition software (e.g., OpenCV)

[1158] Database software (e.g. MongoDB)

[1159] Cloud infrastructure (e.g., Amazon Web Services, AWS, etc.)

[1160] Data collection and analysis

[1161] 1. Collecting language data

[1162] The device's microphone captures surrounding sounds in real time and temporarily stores the audio data in internal memory. Noise filtering technology is used to clear the audio data and make it available for analysis.

[1163] 2. Image data collection

[1164] The device's camera captures images within its field of view in real time and temporarily stores the image data in its internal memory. Based on the high-resolution camera images, it is possible to identify an individual's face.

[1165] 3. Data Analysis

[1166] The collected voice data is converted into text using voice recognition software, and the content of the conversation and keywords are extracted. The image data is analyzed using image recognition software to identify individuals. The results of these analyses are temporarily stored on the device.

[1167] Data transmission and storage

[1168] 1. Data transmission

[1169] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[1170] 2. Data storage

[1171] The analysis data received by the server is stored in a database for each user, with metadata such as date and time, geographical information, and keywords added, and an index is automatically generated.

[1172] Finding and Viewing Data

[1173] 1. Search by keyword

[1174] Users enter keywords they want to search for through a dedicated application or web portal. Keywords refer to specific conversation content or personal information. The server searches for relevant data from the user's database based on the entered keywords, and sends and displays the search results on the user's device.

[1175] 2. Response when you return

[1176] When a customer returns, the device has the ability to search and display previous interactions, allowing store staff to respond quickly based on past customer information.

[1177] Customized product information suggestions

[1178] This includes a means to suggest product information and services that reflect the customer's interaction history and preferences based on the analysis results. For example, if a customer has previously talked about a specific product, other related products can be suggested.

[1179] Specific examples

[1180] Consider an example of use in a cafe. When a customer visits the cafe for the first time and orders a cappuccino, the order details are recorded and analyzed by the terminal and sent to the server. When the customer returns later, the staff member can quickly search for the customer's information and smoothly re-offer the same by asking, "Would you like your usual cappuccino?"

[1181] Prompt Sentence Examples

[1182] "Design a system for a customer support application that searches for a customer's order history based on their previous visits and provides them promptly."

[1183] As a result, the present invention makes it possible to improve the efficiency of customer service and provide customized services.

[1184] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1185] Step 1:

[1186] The device (a smartphone or smart glasses) captures the surrounding audio and video in real time.

[1187] Input: Audio data collected by a microphone, image data collected by a camera.

[1188] Operation: The device temporarily stores audio data in its internal memory and performs noise filtering to obtain clear audio data. Image data is also temporarily stored in its internal memory.

[1189] Output: Noise-filtered audio data and temporarily stored image data.

[1190] Step 2:

[1191] The voice data collected by the device is analyzed using voice recognition software (e.g., Google Cloud Speech-to-Text) and converted into text data.

[1192] Input: Noise-filtered audio data.

[1193] How it works: Speech recognition software converts voice data into text and extracts dialogue and keywords.

[1194] Output: Parsed text data and extracted keywords.

[1195] Step 3:

[1196] The image data collected by the device is analyzed using image recognition software (e.g., OpenCV) to identify individuals.

[1197] Input: Temporarily saved image data.

[1198] How it works: Image recognition software analyzes image data, identifies individuals' faces, and checks them against an existing database of people to confirm a match.

[1199] Output: Result data that identifies the individual.

[1200] Step 4:

[1201] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[1202] Input: Parsed text data, extracted keywords, and result data of personal identification.

[1203] How it works: The device combines all this data into a single package, encrypts it, and sends it to the server.

[1204] Output: Encrypted data package.

[1205] Step 5:

[1206] The server analyzes the received data package and stores the information in a database for each user.

[1207] Input: Encrypted data package.

[1208] How it works: The server decrypts the data package and stores the analyzed language data and image data in a database for each user. When saving, it adds metadata such as date, time, geographical information, and keywords.

[1209] Output: Parsed data stored in a per-user database.

[1210] Step 6:

[1211] Users enter search keywords through a dedicated application or web portal, and the server searches for relevant data.

[1212] Input: The search keyword entered by the user.

[1213] How it works: The server searches for relevant data based on the search keywords in the user's database and filters it using metadata.

[1214] Output: The relevant data found.

[1215] Step 7:

[1216] The server sends the search results to the user's terminal, which displays the results.

[1217] Input: The relevant data found.

[1218] Operation: The server sends the search results to the user's device, which receives and displays the data.

[1219] Output: Search results displayed on the user's device.

[1220] Step 8:

[1221] When a customer visits the store again, the device will recognize the visit and search for and display the previous interaction.

[1222] Input: Personal information of returning customers.

[1223] How it works: The server searches a database of past interactions and sends them to the user's device, which then displays the information.

[1224] Output: Returning customer information and previous interactions displayed on the terminal.

[1225] Step 9:

[1226] Recommend appropriate product information and services based on the customer's interaction history.

[1227] Input: Customer interaction history.

[1228] How it works: The server analyzes past interaction history, compiles relevant product and service information, and sends it to the user's device. The device then displays the information and makes appropriate suggestions.

[1229] Output: Product information and service suggestions displayed on the device.

[1230] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1231] This invention is realized as a system that combines a system that can collect, analyze, store, and search language and image data with an emotion engine that recognizes user emotions. This system consists of a terminal equipped with speech recognition, image recognition, and emotion recognition functions, a server that stores data, and an application that allows users to access the data through an interface.

[1232] Audio and image collection

[1233] Language data collection

[1234] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology.

[1235] Image data collection

[1236] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, which is also temporarily stored in its internal memory. The camera takes high-resolution images and is capable of identifying people's faces.

[1237] emotion recognition

[1238] Emotion Recognition from Speech Data

[1239] The device sends the collected voice data to an emotion recognition engine, which analyzes the user's voice tone, pitch, rhythm, etc. to recognize emotions. The analysis results are saved along with the content of the conversation.

[1240] Emotion recognition from image data

[1241] The device sends the collected image data to an emotion recognition engine, which then analyzes the user's facial expressions and movements to recognize their emotions. The analysis results are then saved along with the person's identification information.

[1242] Data analysis

[1243] Language Data Analysis

[1244] Voice recognition software is used to analyze the voice data collected by the device. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are converted into text format and stored along with emotion data.

[1245] Image data analysis

[1246] Image recognition software is used to analyze the image data collected by the device. The software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also stored along with the emotion data.

[1247] Data transmission and storage

[1248] Sending data

[1249] The terminal packages the analyzed language data, image data, and emotion data and transmits them to the server in an encrypted format.

[1250] Data storage

[1251] The server stores the analyzed data it receives in a database for each user. The database is provided with metadata such as date and time, geographical information, emotional data, and keywords. The database is also designed to automatically generate indexes to improve search efficiency.

[1252] Finding and Viewing Data

[1253] Search by keyword

[1254] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[1255] Searching for Data

[1256] The server searches the user's database for relevant data based on the entered keywords, filtering the data based on date, time, geographical information, emotional data, and indexes.

[1257] Displaying search results

[1258] The server sends the search results to the user's device, which then displays the received search results, allowing the user to visually check the relevant conversation content, the other party's information, and their emotional state.

[1259] Specific examples

[1260] Example 1: Reviewing the conversation and emotions after a business meeting

[1261] The device collects and analyzes voice, image data, and emotional data during a business meeting in real time. After the meeting is over, when the user searches for the keyword "business meeting" in the application, the server finds and displays the relevant analysis data. The user can easily review the details of the meeting, information about the other party, and their own and the other party's emotional states.

[1262] Example 2: When you want to remember the name or feelings of someone you met at an event

[1263] During the event, the device collects and analyzes audio, image data, and emotional data. After the event is over, when the user searches for "event," the server finds relevant data and displays photos, names, conversations, and emotional states. This allows the user to easily recall the names, conversations, and emotional states of people they met at the event.

[1264] The above is an embodiment of a system based on the present invention that includes an emotion engine. This system records many conversations and interactions, manages emotional states, and provides an environment in which necessary information can be easily searched and viewed.

[1265] The processing flow will be explained below.

[1266] Step 1:

[1267] The device collects the audio

[1268] The device's built-in microphone captures surrounding sounds in real time, and the captured audio data is temporarily stored in digital format in the device's internal memory.

[1269] Step 2:

[1270] The device collects the images

[1271] The camera on the device captures images within its field of view in real time, and the captured image data is temporarily stored in the internal memory frame by frame.

[1272] Step 3:

[1273] The device analyzes the audio data

[1274] The device's internal voice recognition software analyzes the temporarily stored voice data to extract the conversation content and keywords. The extracted information is converted into text format, and the analysis results are temporarily stored on the device.

[1275] Step 4:

[1276] The device analyzes the image data

[1277] The device's internal image recognition software analyzes the temporarily stored image data to detect faces. The detected faces are then compared with an existing database to identify specific people. The results of this analysis are also temporarily stored on the device.

[1278] Step 5:

[1279] The device recognizes emotions from voice data

[1280] The device's emotion engine analyzes the voice data and recognizes the user's emotions by analyzing their tone, pitch, rhythm, etc. The emotion analysis results are also temporarily stored on the device.

[1281] Step 6:

[1282] The device recognizes emotions from image data

[1283] The device's emotion engine analyzes the image data and recognizes the user's emotions by analyzing their facial expressions and movements. The results of this emotion analysis are also temporarily stored on the device.

[1284] Step 7:

[1285] The device sends the data to the server

[1286] The results of the voice and image analysis and emotional data are packaged and sent to a server in an encrypted format, along with metadata such as date, time, and location information.

[1287] Step 8:

[1288] The server receives the data

[1289] The server receives the data sent from the device and stores it in a database for each user, where it is indexed for efficient searching.

[1290] Step 9:

[1291] The user enters a keyword

[1292] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[1293] Step 10:

[1294] The server searches for the data

[1295] The server searches the user's database to find data related to the entered keywords, and a search algorithm takes into account date, time, location, index, and sentiment data to filter out relevant data.

[1296] Step 11:

[1297] The server sends the search results

[1298] The filtered search results are sent to the user's device, and are provided in text, image, and audio formats.

[1299] Step 12:

[1300] Your device will display the search results.

[1301] The search results received by the user's device are displayed on the screen, allowing the user to visually check the relevant conversation content, information about the other party, and their emotional state.

[1302] The above are the specific processing steps of the system including the emotion engine.

[1303] Example 2

[1304] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1305] Conventional data collection and analysis systems lack emotion recognition capabilities, making it difficult to accurately grasp the emotional state of users and targets. Furthermore, when storing and searching large amounts of data, the lack of appropriate indexes and metadata makes it difficult to efficiently search and browse information. As a result, users are unable to easily review the content of conversations and related information, potentially leading to delays in productivity and decision-making.

[1306] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting language data, means for collecting image data, means for clearing the collected language data using noise filtering technology, means for capturing the collected image data at high resolution to identify a person, means for transmitting the language data to an emotion recognition engine to analyze voice tone, pitch, rhythm, etc., means for transmitting the image data to the emotion recognition engine to analyze facial expressions, slight movements, etc., means for analyzing the collected language data to extract dialogue content, means for analyzing the collected image data to identify a person, means for transmitting the analyzed language data and image data to the server, means for storing the transmitted data in a database for each user, means for generating an index by adding metadata such as date and time, geographic information, and emotion data to the stored data, means for searching the stored data based on keywords, and means for displaying search results on the user's terminal. This enables effective collection, analysis, and search of dialogue content and related information, including the user's emotional state.

[1307] "Language data" is data that includes the content of communication in audio or text format.

[1308] "Image data" refers to data including still images and moving images collected using a camera or other imaging device.

[1309] "Noise filtering technology" is a technology that removes background noise and unnecessary sounds from collected audio data.

[1310] "High resolution" refers to a resolution that allows images and videos to be displayed clearly down to the fine details.

[1311] An "emotion recognition engine" refers to software or hardware functionality that analyzes audio and image data to identify a user's emotional state.

[1312] "Dialogue content" refers to the content and meaning of the conversation extracted from audio data or text data.

[1313] "Identifying a person" means identifying the face or features of a specific individual from image data.

[1314] "Collected Data" refers to all data collected by the system, including language data and image data.

[1315] A "database" refers to a collection of data organized and stored for a specific purpose.

[1316] "Metadata" refers to data that contains information about the stored data, such as date and time, geographical information, and emotional data.

[1317] An "index" refers to a list or catalog created within a database to make data search and management more efficient.

[1318] A "keyword" refers to a specific word or string of characters that a user enters for searching or filtering.

[1319] This invention is a system that not only collects language data and image data, but also analyzes, stores, and searches that data, but also combines it with an emotion engine that recognizes user emotions. This system is composed of a terminal equipped with speech recognition, image recognition, and emotion recognition functions, a server that stores data, and an application that allows users to access the data through an interface. Specific embodiments are described below.

[1320] System Configuration

[1321] Audio and image collection

[1322] 1. Collecting language data

[1323] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology (e.g., Adobe Audition).

[1324] 2. Image data collection

[1325] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, takes high-resolution photos, and temporarily stores them in its internal memory. The camera has the ability to identify faces, which can be used to identify targets.

[1326] emotion recognition

[1327] 3. Emotion Recognition from Speech Data

[1328] The device sends the collected voice data to an emotion recognition engine (e.g., Affectiva SDK), which analyzes the voice tone, pitch, rhythm, etc. to recognize emotions. The analysis results are saved along with the content of the conversation.

[1329] 4. Emotion Recognition from Image Data

[1330] The image data collected by the device is sent to an emotion recognition engine, which analyzes the user's facial expressions and movements to recognize their emotions. The analysis results are saved along with the person's identification information.

[1331] Data analysis

[1332] 5. Language Data Analysis

[1333] The device uses voice recognition software (e.g., Google Speech-to-Text) to analyze the voice data collected. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are converted into text format and stored along with emotion data.

[1334] 6. Analysis of image data

[1335] Image recognition software (e.g., Amazon Rekognition) is used to analyze the image data collected by the device. This software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also stored along with the emotion data.

[1336] Data transmission and storage

[1337] 7. Data transmission

[1338] The terminal transmits the analyzed language data, image data, and emotion data to the server in an encrypted format.

[1339] 8. Data Retention

[1340] The server stores the analyzed data it receives in a database for each user. At this time, the database is given metadata such as date and time, geographical information, emotion data, and keywords, and an index is automatically generated based on this information.

[1341] Finding and Viewing Data

[1342] 9. Search by Keyword

[1343] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[1344] 10. Searching for Data

[1345] The server searches the user's database for relevant data based on the entered keywords, filtering the data based on date, time, geographical information, emotional data, and indexes.

[1346] 11. Display of search results

[1347] The server sends the search results to the user's device, which then displays the received search results, allowing the user to visually confirm the relevant conversation content, the other party's information, and their emotional state.

[1348] Specific examples

[1349] Example 1: Reviewing the conversation and emotions after a business meeting

[1350] The device collects and analyzes voice, image data, and emotional data during negotiations in real time. After the negotiation is over, when the user searches for the keyword "negotiation" in the application, the server finds and displays the relevant analysis data. The user can easily review details of the negotiation, information about the other party, and even their own and the other party's emotional states.

[1351] Example 2: When you want to remember the name or feelings of someone you met at an event

[1352] During the event, the device collects and analyzes audio, image data, and emotional data. After the event is over, when the user searches for "event," the server finds relevant data and displays photos, names, conversations, and emotional states. This allows the user to easily recall the names, conversations, and emotional states of people they met at the event.

[1353] An example of a prompt to actually enter

[1354] 1. "I want to review the conversation and emotional state of this business meeting."

[1355] 2. "Tell me how you would use this system to remember information about people you met at events."

[1356] This allows the system to effectively collect, analyze, store, search, and display dialogue content and related information for users, improving the user experience.

[1357] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1358] Step 1:

[1359] Language data collection

[1360] The device uses a high-performance microphone to capture surrounding audio. The input is ambient audio data, which the device collects in real time and temporarily stores in its internal memory. Then, noise filtering technology is used to remove background noise and unwanted sounds from the collected audio data. The output is clear audio data.

[1361] Step 2:

[1362] Image data collection

[1363] The device uses a camera to capture video within its field of view in real time. The input is video data of the surroundings, which the device temporarily stores in its internal memory. The image data is captured at high resolution and processed using a function to identify human faces. The output is the image data after face identification.

[1364] Step 3:

[1365] Emotion Recognition from Speech Data

[1366] The device sends the collected voice data to the emotion recognition engine. The input is clear voice data, which the emotion recognition engine analyzes and recognizes the user's emotional state by analyzing the voice tone, pitch, and rhythm. The output is the emotion analysis result.

[1367] Step 4:

[1368] Emotion recognition from image data

[1369] The device sends the collected image data to the emotion recognition engine. The input is image data after face identification, which the emotion recognition engine analyzes to recognize the user's emotional state by analyzing facial expressions and micro-movements. The output is the image emotion analysis results.

[1370] Step 5:

[1371] Language Data Analysis

[1372] The device uses voice recognition software to analyze the collected voice data. The input is clear voice data, which the voice recognition software converts into text. The software then extracts important dialogue and keywords from the text data. The output is the analyzed text data and keywords.

[1373] Step 6:

[1374] Image data analysis

[1375] Image recognition software is used to analyze the image data collected by the device. The input is image data after face identification, which the image recognition software analyzes to detect human faces and identify people by comparing them with information in an existing database. The output is the person recognition result.

[1376] Step 7:

[1377] Sending data

[1378] The terminal packages the analyzed language data, image data, and emotion data and transmits them in encrypted form to the server. The input is the analyzed language data, image data, and emotion data, which the terminal encrypts and transmits, and the output is the encrypted data transmitted to the server.

[1379] Step 8:

[1380] Data storage

[1381] The server stores the received data in a database for each user. The input is encrypted data sent to the server, which decrypts it and stores it in the database. Here, metadata such as date and time, geographic information, emotion data, and keywords are added, and an index is generated. The output is an indexed database entry.

[1382] Step 9:

[1383] Search by keyword

[1384] A user enters search keywords through a dedicated application or web portal. The input is the user-specified search keywords, which are sent to the server. The output is a search query based on the keywords.

[1385] Step 10:

[1386] Searching for Data

[1387] The server searches for relevant data from the user's database based on keywords. The input is a search query, and the server filters the relevant data based on this using date, time, geographic information, sentiment data, and indexes. The output is a dataset of search results.

[1388] Step 11:

[1389] Displaying search results

[1390] The server sends the search results to the user's device. The input is a dataset of search results, which is sent to the user's device. The user's device displays the search results based on the received data, allowing the user to visually confirm related conversation content, information about the other party, and their emotional state. The output is the displayed search results.

[1391] (Application example 2)

[1392] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1393] In traditional brick-and-mortar stores, it was difficult for store associates to understand customer interactions and their emotional state in real time, making it difficult to optimize the quality of service. Furthermore, there was a lack of effective means to make optimal suggestions and follow-ups based on customer emotions, which could result in lower customer satisfaction. Furthermore, traditional systems lacked the ability to provide real-time feedback on collected data, making it difficult for store associates to respond immediately.

[1394] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting language data, means for collecting image data, means for extracting dialogue content from the collected language data, means for identifying people from the collected image data, means for recognizing emotions from the collected language data and image data, means for transmitting the analyzed language data, image data, and emotion data to the server, means for storing the transmitted data in a database for each user, means for providing feedback on the stored data in real time, means for searching the stored data based on keywords, and means for displaying search results on the user's terminal. This allows store clerks to grasp the dialogue and emotional state of customers in real time and respond appropriately, improving the quality of service and increasing customer satisfaction.

[1395] The "means for collecting language data" refers to a device or method that uses a microphone or the like to capture surrounding sounds and collect them as language data.

[1396] "Means for collecting image data" refers to a device or method that uses a camera or the like to capture images within the field of view and collects them as image data.

[1397] The "means for extracting dialogue content from collected language data" refers to a device or method that converts collected language data into text using speech recognition technology and extracts important dialogue content and keywords.

[1398] "Means for identifying a person from collected image data" refers to a device or method that uses image recognition technology to identify a person from collected image data and compare it with a database.

[1399] The "means for recognizing emotions from collected language data and image data" refers to a device or method for recognizing the emotional state of a user by analyzing voice tone, facial expressions, etc.

[1400] The "means for transmitting analyzed language data, image data, and emotion data to a server" refers to a device or method for packaging the analyzed data and transmitting it to a server via a network.

[1401] The "means for storing transmitted data in a database for each user" refers to a device or method for sorting and storing the analysis data received by the server for each user.

[1402] "Means for providing real-time feedback of stored data" refers to a device or method that immediately processes the analyzed and stored data and displays it in real-time through a user interface.

[1403] The term "means for searching stored data based on keywords" refers to a device or method for efficiently searching related information from a stored database based on keywords entered by a user.

[1404] "Means for displaying search results on the user's terminal" refers to a device or method for transmitting searched data to the user's terminal and visually displaying it through an interface.

[1405] This invention is a system that enables store clerks in brick-and-mortar stores to grasp the conversations and emotional state of customers in real time and optimize the quality of service. The system consists of smart glasses, a voice recognition engine, an image recognition engine, an emotion recognition engine, a server, cloud storage, and a real-time display. This system allows the details of conversations with customers and the emotional state of customers to be grasped instantly, enabling the provision of appropriate services.

[1406] Hardware and Software Configuration

[1407] Smart Glasses

[1408] The smart glasses are equipped with high-performance microphones and cameras that collect voice and image data from customer interactions in real time.

[1409] Speech Recognition Engine

[1410] The speech recognition engine analyzes the collected voice data and converts it into text. This analysis is performed using, for example, the Google Speech Recognition API.

[1411] Image Recognition Engine

[1412] The image recognition engine analyzes the collected image data and identifies the customer's face using tools such as OpenCV and DeepFace.

[1413] Emotion Recognition Engine

[1414] The emotion recognition engine analyzes collected voice and image data to recognize the customer's emotional state, using the Emotion API and machine learning models for emotion recognition.

[1415] Server and Cloud Storage

[1416] The collected data and analysis results are sent in encrypted form to a server and stored in a per-user database using cloud storage services such as AWS or Google Cloud Storage.

[1417] Real-time Display

[1418] Store associates receive real-time feedback through the smart glasses' display, allowing them to respond immediately to the customer's immediate needs and emotional state.

[1419] Processing flow explanation

[1420] 1. Audio and Image Collection:

[1421] The smart glasses collect customer interaction through a microphone and capture images of the surroundings through a camera, and these data are temporarily stored in the internal memory.

[1422] 2. Data Analysis:

[1423] The voice data is converted into text by a voice recognition engine, which then analyzes emotions from the tone and pitch of the voice, while the image data is analyzed by an image recognition engine to identify the customer's face and recognize their emotional state from their facial expressions.

[1424] 3. Data transmission and storage:

[1425] The analyzed data is encrypted and sent to a server, which then stores the data in cloud storage and organizes it for each user.

[1426] 4. Real-time feedback:

[1427] The smart glasses display allows store staff to see the customer's conversation and emotional state in real time, allowing them to provide optimal responses and suggestions based on this feedback.

[1428] Specific examples

[1429] 1. Know when a customer is interested in a specific product:

[1430] The program recognizes the customer's interests in real time and provides feedback to the store clerk, saying, "You are interested in this product."

[1431] 2. Recognize when a customer is unhappy:

[1432] The program detects a customer's dissatisfied facial expression and displays a warning to the store clerk saying, "The customer is dissatisfied."

[1433] Prompt Sentence Examples

[1434] An example of a prompt is: "When a customer asks a question about a particular product, the smart glasses analyze the conversation and the customer's facial expressions in real time and display feedback such as, 'It seems you are interested in this product. Further explanation is needed.'"

[1435] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1436] Step 1:

[1437] The device (smart glasses) collects the conversational voice with the customer using a microphone and captures the surrounding field of view using a camera. These data are temporarily stored in the internal memory. Voice data and image data are collected as input and saved in the internal memory as output. In concrete terms, the microphone captures the voice signal and the camera continuously takes images.

[1438] Step 2:

[1439] The device sends the collected voice data to a speech recognition engine, which converts it into text data. The input is voice data, which the speech recognition engine converts into text. Text data is generated as output. Specifically, the device uses the Google Speech Recognition API to convert the voice signal into text.

[1440] Step 3:

[1441] The device sends the collected image data to an image recognition engine to identify the customer's face. The input is image data, which the image recognition engine analyzes to identify the person. The output is information about the identified person. Specifically, the device uses OpenCV and DeepFace to analyze facial features and compare them with a database.

[1442] Step 4:

[1443] The device sends the collected voice and image data to the emotion recognition engine to analyze the customer's emotional state. The input is voice and image data, which the emotion recognition engine analyzes to obtain emotional data. Emotion data is generated as output. Specifically, the device uses the Emotion API to analyze voice tone and facial expressions.

[1444] Step 5:

[1445] The device encrypts the analyzed language data, image data, and emotion data and sends it to the server. The input is the analyzed data, which is encrypted and sent to the server via the network. The output is the data sent to the server. Specifically, the device sends the data using an encryption protocol such as SSL / TLS.

[1446] Step 6:

[1447] The server stores the analysis data it receives in a database for each user. The input is the analyzed data, which is organized for each user and stored in the database. The output is the analysis data stored in the database. Specifically, the server organizes and stores the data using AWS or Google Cloud Storage.

[1448] Step 7:

[1449] The server sends analysis data for feedback to the device in real time. The input is analysis data obtained from the user's database, which is sent to the device in real time. The output is feedback that is displayed immediately on the device. Specifically, the server sends data immediately using a protocol such as WebSocket.

[1450] Step 8:

[1451] Based on the feedback displayed on the terminal in real time, the user (store clerk) responds appropriately to the customer. The input is feedback data from the terminal, and the user provides services based on this. The output is the result of providing an appropriate service. In concrete terms, the user checks the display on the smart glasses and engages in a dialogue based on the feedback content.

[1452] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1453] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1454] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1455] [Fourth embodiment]

[1456] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1457] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1458] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1459] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1460] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1461] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1462] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1463] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1464] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1465] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1466] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1467] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1468] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1469] The present invention is realized as a system that can collect, analyze, store, and search language and image data. This system consists of a terminal with speech recognition and image recognition capabilities, a server that stores the data, and an application that allows users to access the data through an interface.

[1470] Audio and image collection

[1471] Language data collection

[1472] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology.

[1473] Image data collection

[1474] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, which is also temporarily stored in its internal memory. The camera takes high-resolution images and is capable of identifying people's faces.

[1475] Data analysis

[1476] Language Data Analysis

[1477] Voice recognition software is used to analyze the voice data collected by the device. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are temporarily stored on the device.

[1478] Image data analysis

[1479] Image recognition software is used to analyze the image data collected by the device. This software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also temporarily stored within the device.

[1480] Data transmission and storage

[1481] Sending data

[1482] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[1483] Data storage

[1484] The server stores the analysis data it receives in a database for each user. The database is provided with metadata such as date and time, geographical information, and keywords. The database is also designed to automatically generate indexes to improve search efficiency.

[1485] Finding and Viewing Data

[1486] Search by keyword

[1487] Users enter keywords they want to search for through a dedicated application or web portal, which refer to specific conversation content or personal information.

[1488] Searching for Data

[1489] The server searches the user's database for relevant data based on the keywords entered, filtering the data based on metadata such as date, time, and geographic location.

[1490] Displaying search results

[1491] The server sends the search results to the user's device, which displays the received search results, allowing the user to check past conversations and personal information.

[1492] Specific examples

[1493] Example 1: When you want to review the conversation after a business meeting

[1494] The device collects and analyzes audio and video during business negotiations in real time. After the negotiation is over, when the user searches for the keyword "business negotiation" in the application, the server finds and displays the relevant analysis data. The user can easily review the details of the negotiation and the other party's information.

[1495] Example 2: You want to remember the name of someone you met at an event

[1496] During the event, the device collects and analyzes audio and images. After the event is over, when a user searches for "event," the server finds relevant data and displays photos, names, conversations, etc. This allows users to easily recall the names and conversations of people they met at the event.

[1497] The above is an embodiment of the system according to the present invention. This system records many conversations and contacts and provides an environment in which necessary information can be easily searched and viewed.

[1498] The processing flow will be explained below.

[1499] Step 1:

[1500] The device collects the audio

[1501] The device's built-in microphone captures surrounding sounds in real time, and the captured audio data is temporarily stored in digital format in the device's internal memory.

[1502] Step 2:

[1503] The device collects the images

[1504] The camera on the device captures images within its field of view in real time, and the captured image data is temporarily stored in the internal memory frame by frame.

[1505] Step 3:

[1506] The device analyzes the audio data

[1507] The device's internal voice recognition software analyzes the temporarily stored voice data to extract the conversation content and keywords. The extracted information is converted into text format, and the analysis results are temporarily stored on the device.

[1508] Step 4:

[1509] The device analyzes the image data

[1510] The device's internal image recognition software analyzes the temporarily stored image data to detect faces. The detected faces are then compared with an existing database to identify specific people. The results of this analysis are also temporarily stored on the device.

[1511] Step 5:

[1512] The device sends the data to the server

[1513] The results of the audio and image analysis are packaged and sent to a server in an encrypted format, along with metadata such as date, time, and location information.

[1514] Step 6:

[1515] The server receives the data

[1516] The server receives the data sent from the device and stores it in a database for each user, where it is indexed for efficient searching.

[1517] Step 7:

[1518] The user enters a keyword

[1519] Users enter keywords for the content they want to search for through a dedicated application or web portal. Keywords refer to specific conversation content or personal information.

[1520] Step 8:

[1521] The server searches for the data

[1522] The server searches the user's database for data related to the entered keywords, and a search algorithm takes into account date, time, location, and indexing to filter out relevant data.

[1523] Step 9:

[1524] The server sends the search results

[1525] The filtered search results are sent to the user's device, and are provided in text, image, and audio formats.

[1526] Step 10:

[1527] Your device will display the search results.

[1528] The search results received by the user's device are displayed on the screen, allowing the user to visually check related conversation content and other party information.

[1529] The above are the specific processing steps of the system.

[1530] Example 1

[1531] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1532] Conventional data collection and analysis systems are inefficient in collecting, analyzing, storing, and searching language and image data, making it difficult for users to quickly obtain the information they need. Furthermore, the lack of quality improvement measures, such as noise filtering for voice data and high-resolution image capture, reduces the reliability of data and reduces search efficiency.

[1533] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1534] In this invention, the server includes means for collecting language data, means for collecting image data, means for analyzing the collected language data to extract important information and keywords, means for analyzing the collected image data to identify people, means for packaging the analyzed language data and image data, encrypting the data, and transmitting the packaged data to the server, means for storing the transmitted data in a database for each user and adding metadata, means for searching indexed data based on keywords, means for displaying search results on the user's device, means for applying noise filtering technology to clarify voice data, means for capturing image data at high resolution and performing facial recognition, means for a user to input search keywords through an application or web portal, and for the server to efficiently search for related data based on the index, and means for visually displaying search results, thereby improving the quality of collected data and enabling users to quickly search for and confirm the information they need.

[1535] "Means for collecting language data" refers to devices or software for capturing and storing speech as digital data.

[1536] "Means for collecting image data" refers to a device or software for capturing and storing visual information as a digital image.

[1537] "Means of analyzing collected language data to extract key information and keywords" refers to software or algorithms that convert speech data into text data and identify specific words or phrases within it.

[1538] "Means for analyzing collected image data to identify persons" refers to software or algorithms that recognize human faces in images and identify specific persons by matching them with existing databases.

[1539] "Means for packaging analyzed language data and image data, encrypting them, and transmitting them to a server" refers to a device or software for packaging the analyzed information into a single data package, encrypting it, and transmitting it to a server via a communications network.

[1540] "Means for storing the transmitted data in a database for each user and adding metadata" refers to a device or software for classifying data based on the identification information of each user, adding supplementary information such as date and time and geographical information, and storing the data in a database.

[1541] "Means for searching indexed data based on keywords" refers to algorithms or software for efficiently searching data within a database using specific keywords.

[1542] "Means for displaying search results on a user's terminal" refers to a device or software for visually displaying the search result data sent from the server on a user interface.

[1543] "Means for applying noise filtering techniques to make the audio data clear" refers to algorithms or software that remove background noise from the audio data to improve the sound quality.

[1544] "Means for capturing high-resolution image data and performing facial recognition" refers to a device or software that captures high-resolution photographs and identifies human faces from those images.

[1545] "Means for a user to input search keywords through an application or web portal and for the server to efficiently search for related data based on an index" refers to a device or software that allows a user to input keywords through an interface and for the server to quickly search for data related to those keywords.

[1546] "Means for visually displaying search results" refers to a device or software that visually displays the search results provided by the server on a screen so that the user can confirm the results.

[1547] This invention is a system that collects, analyzes, stores, and searches voice and image data. This system consists of a terminal equipped with voice recognition and image recognition functions, a server for storing data, and an application for user access.

[1548] Audio data collection

[1549] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. For example, this microphone is used to obtain high-quality recordings of audio during meetings. The collected audio data is temporarily stored in the device's internal memory. Noise filtering technology (e.g., noise cancellation software) is used to improve the quality of the audio data.

[1550] Image data collection

[1551] The device is equipped with a high-resolution camera that can capture images of what is within its field of view in real time. For example, this camera can be used to capture images of meeting scenes and participants' faces. The collected image data is temporarily stored in internal memory. Furthermore, facial recognition software can be used to identify people's faces from the captured images.

[1552] Data analysis

[1553] Dedicated analysis software is used to analyze the voice and image data collected by the device. For voice data, voice recognition software (e.g., Google Speech-to-Text API) is used to convert the data into text and extract important dialogue and keywords. For image data, image recognition software (e.g., OpenCV) is used to detect people's faces and match them with existing information in a database. The analysis results are temporarily stored inside the device.

[1554] Data transmission and storage

[1555] The device packages the analyzed language data and image data, encrypts them, and sends them to the server. The data is encrypted using technology such as AES encryption. The server stores the received data in a database for each user, adding metadata such as date and time, geographic information, and keywords. It also automatically generates an index to improve search efficiency.

[1556] Searching for Data

[1557] Users can search for data by entering specific keywords through a dedicated application or web portal. The server then references the index based on the entered keywords and efficiently searches for related data. For example, users can enter keywords such as "business negotiation" or "event" to find related analysis data.

[1558] Displaying search results

[1559] The server sends the search results to the user's device, which then visually displays the results, allowing the user to check past interactions and personal information.

[1560] Specific examples

[1561] Example 1: When you want to review the conversation after a business meeting

[1562] The device collects and analyzes audio and video during the negotiation in real time. After the negotiation is over, when the user searches for the keyword "negotiation" in the application, the server finds the relevant analysis data and displays the results. This allows the user to easily review the details of the negotiation and the other party's information.

[1563] Example 2: You want to remember the name of someone you met at an event

[1564] During the event, the device collects and analyzes audio and images. After the event ends, when the user searches for "event" in the application, the server finds relevant data and displays photos, names, conversation details, etc. This allows the user to easily remember the names and conversation details of people they met at the event.

[1565] Prompt Sentence Examples

[1566] "I want to review the content of yesterday's business meeting."

[1567] "I forgot the name of the person I met at the event last week."

[1568] This concludes the description of the embodiment of the invention. This system improves the quality of data and enables users to quickly search and check the information they need.

[1569] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1570] Step 1: Collecting language and image data

[1571] The device activates a high-performance microphone and a high-resolution camera to capture the surrounding audio and video in real time. For example, it collects audio and video of participants in a meeting. This is the input, and the data is temporarily stored in the internal memory. The output is raw audio and video data.

[1572] Specific behavior:

[1573] The device activates the microphone and captures audio data.

[1574] The device activates the camera and captures image data.

[1575] The collected audio and image data is temporarily stored in the internal memory.

[1576] Step 2: Filtering the data

[1577] The collected data contains noise and unnecessary information, so the device applies noise filtering technology. The audio data is filtered to make it clear, and unnecessary parts of the image data are removed. This improves the quality of the data. The input is raw audio data and image data, and the output is clear audio data and high-quality image data.

[1578] Specific behavior:

[1579] The device uses noise cancelling software to remove noise from the audio data.

[1580] The terminal filters out unnecessary parts of the image data.

[1581] Step 3: Analyze the data

[1582] The device runs analysis software to convert the filtered voice data into text and extract key keywords and dialogue. Similarly, facial recognition software is used on image data to identify people. The input is clear voice data and high-quality image data, and the output is analyzed text data and facial recognition results.

[1583] Specific behavior:

[1584] The device converts the voice data into text using the Google Speech-to-Text API.

[1585] The device extracts important keywords from the text.

[1586] The device uses OpenCV to identify people from image data.

[1587] Step 4: Sending data

[1588] The device packages the analyzed language data and image data, encrypts them using AES encryption, and then sends them to the server. The input is the analyzed text data and face recognition results, and the output is an encrypted data packet.

[1589] Specific behavior:

[1590] The device packages the analytical data together.

[1591] The device encrypts the data using AES encryption.

[1592] The device sends the encrypted data to the server.

[1593] Step 5: Save your data

[1594] The server decrypts the received encrypted data and stores it in a database for each user. Metadata such as date, time, geographical information, and keywords are added to the analyzed data. An index is also generated to enable efficient data searches. The input is the encrypted data packet, and the output is the analyzed data stored in the database.

[1595] Specific behavior:

[1596] The server decrypts the encrypted data.

[1597] The server adds metadata to the data.

[1598] The server stores the analysis data in a database.

[1599] The server generates the index.

[1600] Step 6: Search for data

[1601] A user enters search keywords through a dedicated application or web portal. The server references the index and efficiently searches for relevant data. The input is the user's search keywords, and the output is the search result data.

[1602] Specific behavior:

[1603] The user enters search keywords in the application.

[1604] The server references the index to find the relevant data.

[1605] Step 7: Viewing search results

[1606] The server sends the search results to the user's device, which then visually displays the results. The user can review the search results and refer to past interactions and personal information. The input is the search result data, and the output is the visually displayed search results.

[1607] Specific behavior:

[1608] The server sends the search results to the user's terminal.

[1609] The device converts the search results into a display format and displays them to the user in an easy-to-view format.

[1610] (Application example 1)

[1611] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1612] Traditional customer service in brick-and-mortar stores lacks an efficient method for managing customer information when customers return or for referencing past conversations. This results in low accuracy in customer recognition and response, making it difficult to improve customer experience. It also makes it difficult to offer product suggestions or services based on past conversation information, resulting in missed opportunities. Therefore, it is necessary to establish a method for streamlining customer service and providing personalized services.

[1613] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1614] In this invention, the server includes means for collecting language data, means for collecting image data, means for analyzing the collected language data to extract the content of the conversation, means for analyzing the collected image data to identify individuals, means for transmitting the analyzed language data and image data to a remote server, means for storing the transmitted data in a database for each user, means for searching the stored data based on keywords, means for displaying the search results on the user's terminal, means for searching and displaying previous interactions when the user visits again, and means for suggesting product information and services based on the customer's conversation history. This makes it possible to improve the efficiency of customer service and provide personalized services.

[1615] "Language Data" means speech information collected using speech recognition technology.

[1616] "Image data" refers to visual information captured using a camera or other imaging device.

[1617] "Dialogue content" refers to the specific content of the conversation obtained by analyzing the collected language data.

[1618] An "individual" is a specific person who can be identified by analyzing image data.

[1619] A "remote server" is a data storage and processing server accessible over a network.

[1620] A "user-specific database" is a database for storing and managing unique data for each user.

[1621] "Keywords" are important words or phrases that represent specific information and are used to conduct a search.

[1622] "User Terminal" means the electronic device used by a User to access the System.

[1623] "Previous interactions" refers to records of past conversations and responses with customers.

[1624] "Product information" refers to information about the details and specifications of the products being sold.

[1625] "Services" refers to various conveniences and support provided to users.

[1626] The present invention is realized as a system that can collect, analyze, store, and search language and image data. This system is composed of the following elements:

[1627] Hardware and Software

[1628] 1. Device:

[1629] Smartphones (e.g., general smartphones, iPhones, etc.)

[1630] Smart glasses (e.g., Google Glass)

[1631] The device is equipped with a high-performance microphone and camera, allowing it to capture audio and images in real time.

[1632] 2. Software:

[1633] Voice recognition software (e.g., Google Cloud Speech-to-Text)

[1634] Image recognition software (e.g., OpenCV)

[1635] Database software (e.g. MongoDB)

[1636] Cloud infrastructure (e.g., Amazon Web Services, AWS, etc.)

[1637] Data collection and analysis

[1638] 1. Collecting language data

[1639] The device's microphone captures surrounding sounds in real time and temporarily stores the audio data in internal memory. Noise filtering technology is used to clear the audio data and make it available for analysis.

[1640] 2. Image data collection

[1641] The device's camera captures images within its field of view in real time and temporarily stores the image data in its internal memory. Based on the high-resolution camera images, it is possible to identify an individual's face.

[1642] 3. Data Analysis

[1643] The collected voice data is converted into text using voice recognition software, and the content of the conversation and keywords are extracted. The image data is analyzed using image recognition software to identify individuals. The results of these analyses are temporarily stored on the device.

[1644] Data transmission and storage

[1645] 1. Data transmission

[1646] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[1647] 2. Data storage

[1648] The analysis data received by the server is stored in a database for each user, with metadata such as date and time, geographical information, and keywords added, and an index is automatically generated.

[1649] Finding and Viewing Data

[1650] 1. Search by keyword

[1651] Users enter keywords they want to search for through a dedicated application or web portal. Keywords refer to specific conversation content or personal information. The server searches for relevant data from the user's database based on the entered keywords, and sends and displays the search results on the user's device.

[1652] 2. Response when you return

[1653] When a customer returns, the device has the ability to search and display previous interactions, allowing store staff to respond quickly based on past customer information.

[1654] Customized product information suggestions

[1655] This includes a means to suggest product information and services that reflect the customer's interaction history and preferences based on the analysis results. For example, if a customer has previously talked about a specific product, other related products can be suggested.

[1656] Specific examples

[1657] Consider an example of use in a cafe. When a customer visits the cafe for the first time and orders a cappuccino, the order details are recorded and analyzed by the terminal and sent to the server. When the customer returns later, the staff member can quickly search for the customer's information and smoothly re-offer the same by asking, "Would you like your usual cappuccino?"

[1658] Prompt Sentence Examples

[1659] "Design a system for a customer support application that searches for a customer's order history based on their previous visits and provides them promptly."

[1660] As a result, the present invention makes it possible to improve the efficiency of customer service and provide customized services.

[1661] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1662] Step 1:

[1663] The device (a smartphone or smart glasses) captures the surrounding audio and video in real time.

[1664] Input: Audio data collected by a microphone, image data collected by a camera.

[1665] Operation: The device temporarily stores audio data in its internal memory and performs noise filtering to obtain clear audio data. Image data is also temporarily stored in its internal memory.

[1666] Output: Noise-filtered audio data and temporarily stored image data.

[1667] Step 2:

[1668] The voice data collected by the device is analyzed using voice recognition software (e.g., Google Cloud Speech-to-Text) and converted into text data.

[1669] Input: Noise-filtered audio data.

[1670] How it works: Speech recognition software converts voice data into text and extracts dialogue and keywords.

[1671] Output: Parsed text data and extracted keywords.

[1672] Step 3:

[1673] The image data collected by the device is analyzed using image recognition software (e.g., OpenCV) to identify individuals.

[1674] Input: Temporarily saved image data.

[1675] How it works: Image recognition software analyzes image data, identifies individuals' faces, and checks them against an existing database of people to confirm a match.

[1676] Output: Result data that identifies the individual.

[1677] Step 4:

[1678] The terminal packages the analyzed language data and image data, encrypts them, and transmits them to the server.

[1679] Input: Parsed text data, extracted keywords, and result data of personal identification.

[1680] How it works: The device combines all this data into a single package, encrypts it, and sends it to the server.

[1681] Output: Encrypted data package.

[1682] Step 5:

[1683] The server analyzes the received data package and stores the information in a database for each user.

[1684] Input: Encrypted data package.

[1685] How it works: The server decrypts the data package and stores the analyzed language data and image data in a database for each user. When saving, it adds metadata such as date, time, geographical information, and keywords.

[1686] Output: Parsed data stored in a per-user database.

[1687] Step 6:

[1688] Users enter search keywords through a dedicated application or web portal, and the server searches for relevant data.

[1689] Input: The search keyword entered by the user.

[1690] How it works: The server searches for relevant data based on the search keywords in the user's database and filters it using metadata.

[1691] Output: The relevant data found.

[1692] Step 7:

[1693] The server sends the search results to the user's terminal, which displays the results.

[1694] Input: The relevant data found.

[1695] Operation: The server sends the search results to the user's device, which receives and displays the data.

[1696] Output: Search results displayed on the user's device.

[1697] Step 8:

[1698] When a customer visits the store again, the device will recognize the visit and search for and display the previous interaction.

[1699] Input: Personal information of returning customers.

[1700] How it works: The server searches a database of past interactions and sends them to the user's device, which then displays the information.

[1701] Output: Returning customer information and previous interactions displayed on the terminal.

[1702] Step 9:

[1703] Recommend appropriate product information and services based on the customer's interaction history.

[1704] Input: Customer interaction history.

[1705] How it works: The server analyzes past interaction history, compiles relevant product and service information, and sends it to the user's device. The device then displays the information and makes appropriate suggestions.

[1706] Output: Product information and service suggestions displayed on the device.

[1707] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1708] This invention is realized as a system that combines a system that can collect, analyze, store, and search language and image data with an emotion engine that recognizes user emotions. This system consists of a terminal equipped with speech recognition, image recognition, and emotion recognition functions, a server that stores data, and an application that allows users to access the data through an interface.

[1709] Audio and image collection

[1710] Language data collection

[1711] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology.

[1712] Image data collection

[1713] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, which is also temporarily stored in its internal memory. The camera takes high-resolution images and is capable of identifying people's faces.

[1714] emotion recognition

[1715] Emotion Recognition from Speech Data

[1716] The device sends the collected voice data to an emotion recognition engine, which analyzes the user's voice tone, pitch, rhythm, etc. to recognize emotions. The analysis results are saved along with the content of the conversation.

[1717] Emotion recognition from image data

[1718] The device sends the collected image data to an emotion recognition engine, which then analyzes the user's facial expressions and movements to recognize their emotions. The analysis results are then saved along with the person's identification information.

[1719] Data analysis

[1720] Language Data Analysis

[1721] Voice recognition software is used to analyze the voice data collected by the device. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are converted into text format and stored along with emotion data.

[1722] Image data analysis

[1723] Image recognition software is used to analyze the image data collected by the device. The software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also stored along with the emotion data.

[1724] Data transmission and storage

[1725] Sending data

[1726] The terminal packages the analyzed language data, image data, and emotion data and transmits them to the server in an encrypted format.

[1727] Data storage

[1728] The server stores the analyzed data it receives in a database for each user. The database is provided with metadata such as date and time, geographical information, emotional data, and keywords. The database is also designed to automatically generate indexes to improve search efficiency.

[1729] Finding and Viewing Data

[1730] Search by keyword

[1731] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[1732] Searching for Data

[1733] The server searches the user's database for relevant data based on the entered keywords, filtering the data based on date, time, geographical information, emotional data, and indexes.

[1734] Displaying search results

[1735] The server sends the search results to the user's device, which then displays the received search results, allowing the user to visually check the relevant conversation content, the other party's information, and their emotional state.

[1736] Specific examples

[1737] Example 1: Reviewing the conversation and emotions after a business meeting

[1738] The device collects and analyzes voice, image data, and emotional data during a business meeting in real time. After the meeting is over, when the user searches for the keyword "business meeting" in the application, the server finds and displays the relevant analysis data. The user can easily review the details of the meeting, information about the other party, and their own and the other party's emotional states.

[1739] Example 2: When you want to remember the name or feelings of someone you met at an event

[1740] During the event, the device collects and analyzes audio, image data, and emotional data. After the event is over, when the user searches for "event," the server finds relevant data and displays photos, names, conversations, and emotional states. This allows the user to easily recall the names, conversations, and emotional states of people they met at the event.

[1741] The above is an embodiment of a system based on the present invention that includes an emotion engine. This system records many conversations and interactions, manages emotional states, and provides an environment in which necessary information can be easily searched and viewed.

[1742] The processing flow will be explained below.

[1743] Step 1:

[1744] The device collects the audio

[1745] The device's built-in microphone captures surrounding sounds in real time, and the captured audio data is temporarily stored in digital format in the device's internal memory.

[1746] Step 2:

[1747] The device collects the images

[1748] The camera on the device captures images within its field of view in real time, and the captured image data is temporarily stored in the internal memory frame by frame.

[1749] Step 3:

[1750] The device analyzes the audio data

[1751] The device's internal voice recognition software analyzes the temporarily stored voice data to extract the conversation content and keywords. The extracted information is converted into text format, and the analysis results are temporarily stored on the device.

[1752] Step 4:

[1753] The device analyzes the image data

[1754] The device's internal image recognition software analyzes the temporarily stored image data to detect faces. The detected faces are then compared with an existing database to identify specific people. The results of this analysis are also temporarily stored on the device.

[1755] Step 5:

[1756] The device recognizes emotions from voice data

[1757] The device's emotion engine analyzes the voice data and recognizes the user's emotions by analyzing their tone, pitch, rhythm, etc. The emotion analysis results are also temporarily stored on the device.

[1758] Step 6:

[1759] The device recognizes emotions from image data

[1760] The device's emotion engine analyzes the image data and recognizes the user's emotions by analyzing their facial expressions and movements. The results of this emotion analysis are also temporarily stored on the device.

[1761] Step 7:

[1762] The device sends the data to the server

[1763] The results of the voice and image analysis and emotional data are packaged and sent to a server in an encrypted format, along with metadata such as date, time, and location information.

[1764] Step 8:

[1765] The server receives the data

[1766] The server receives the data sent from the device and stores it in a database for each user, where it is indexed for efficient searching.

[1767] Step 9:

[1768] The user enters a keyword

[1769] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[1770] Step 10:

[1771] The server searches for the data

[1772] The server searches the user's database to find data related to the entered keywords, and a search algorithm takes into account date, time, location, index, and sentiment data to filter out relevant data.

[1773] Step 11:

[1774] The server sends the search results

[1775] The filtered search results are sent to the user's device, and are provided in text, image, and audio formats.

[1776] Step 12:

[1777] Your device will display the search results.

[1778] The search results received by the user's device are displayed on the screen, allowing the user to visually check the relevant conversation content, information about the other party, and their emotional state.

[1779] The above are the specific processing steps of the system including the emotion engine.

[1780] Example 2

[1781] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1782] Conventional data collection and analysis systems lack emotion recognition capabilities, making it difficult to accurately grasp the emotional state of users and targets. Furthermore, when storing and searching large amounts of data, the lack of appropriate indexes and metadata makes it difficult to efficiently search and browse information. As a result, users are unable to easily review the content of conversations and related information, potentially leading to delays in productivity and decision-making.

[1783] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting language data, means for collecting image data, means for clearing the collected language data using noise filtering technology, means for capturing the collected image data at high resolution to identify a person, means for transmitting the language data to an emotion recognition engine to analyze voice tone, pitch, rhythm, etc., means for transmitting the image data to the emotion recognition engine to analyze facial expressions, slight movements, etc., means for analyzing the collected language data to extract dialogue content, means for analyzing the collected image data to identify a person, means for transmitting the analyzed language data and image data to the server, means for storing the transmitted data in a database for each user, means for generating an index by adding metadata such as date and time, geographic information, and emotion data to the stored data, means for searching the stored data based on keywords, and means for displaying search results on the user's terminal. This enables effective collection, analysis, and search of dialogue content and related information, including the user's emotional state.

[1784] "Language data" is data that includes the content of communication in audio or text format.

[1785] "Image data" refers to data including still images and moving images collected using a camera or other imaging device.

[1786] "Noise filtering technology" is a technology that removes background noise and unnecessary sounds from collected audio data.

[1787] "High resolution" refers to a resolution that allows images and videos to be displayed clearly down to the fine details.

[1788] An "emotion recognition engine" refers to software or hardware functionality that analyzes audio and image data to identify a user's emotional state.

[1789] "Dialogue content" refers to the content and meaning of the conversation extracted from audio data or text data.

[1790] "Identifying a person" means identifying the face or features of a specific individual from image data.

[1791] "Collected Data" refers to all data collected by the system, including language data and image data.

[1792] A "database" refers to a collection of data organized and stored for a specific purpose.

[1793] "Metadata" refers to data that contains information about the stored data, such as date and time, geographical information, and emotional data.

[1794] An "index" refers to a list or catalog created within a database to make data search and management more efficient.

[1795] A "keyword" refers to a specific word or string of characters that a user enters for searching or filtering.

[1796] This invention is a system that not only collects language data and image data, but also analyzes, stores, and searches that data, but also combines it with an emotion engine that recognizes user emotions. This system is composed of a terminal equipped with speech recognition, image recognition, and emotion recognition functions, a server that stores data, and an application that allows users to access the data through an interface. Specific embodiments are described below.

[1797] System Configuration

[1798] Audio and image collection

[1799] 1. Collecting language data

[1800] The device is equipped with a high-performance microphone that can capture surrounding sounds in real time. The device collects audio data and temporarily stores it in its internal memory. The collected audio data is then cleared using noise filtering technology (e.g., Adobe Audition).

[1801] 2. Image data collection

[1802] The device is equipped with a camera that can capture images within its field of view in real time. The device collects image data, takes high-resolution photos, and temporarily stores them in its internal memory. The camera has the ability to identify faces, which can be used to identify targets.

[1803] emotion recognition

[1804] 3. Emotion Recognition from Speech Data

[1805] The device sends the collected voice data to an emotion recognition engine (e.g., Affectiva SDK), which analyzes the voice tone, pitch, rhythm, etc. to recognize emotions. The analysis results are saved along with the content of the conversation.

[1806] 4. Emotion Recognition from Image Data

[1807] The image data collected by the device is sent to an emotion recognition engine, which analyzes the user's facial expressions and movements to recognize their emotions. The analysis results are saved along with the person's identification information.

[1808] Data analysis

[1809] 5. Language Data Analysis

[1810] The device uses voice recognition software (e.g., Google Speech-to-Text) to analyze the voice data collected. This software converts the voice data into text and extracts important dialogue and keywords. The results of this analysis are converted into text format and stored along with emotion data.

[1811] 6. Analysis of image data

[1812] Image recognition software (e.g., Amazon Rekognition) is used to analyze the image data collected by the device. This software detects human faces and identifies them by matching them with information in an existing database. The results of this analysis are also stored along with the emotion data.

[1813] Data transmission and storage

[1814] 7. Data transmission

[1815] The terminal transmits the analyzed language data, image data, and emotion data to the server in an encrypted format.

[1816] 8. Data Retention

[1817] The server stores the analyzed data it receives in a database for each user. At this time, the database is given metadata such as date and time, geographical information, emotion data, and keywords, and an index is automatically generated based on this information.

[1818] Finding and Viewing Data

[1819] 9. Search by Keyword

[1820] Users enter keywords for their search through a dedicated application or web portal, which can refer to specific conversations, personal information, or emotional states.

[1821] 10. Searching for Data

[1822] The server searches the user's database for relevant data based on the entered keywords, filtering the data based on date, time, geographical information, emotional data, and indexes.

[1823] 11. Display of search results

[1824] The server sends the search results to the user's device, which then displays the received search results, allowing the user to visually confirm the relevant conversation content, the other party's information, and their emotional state.

[1825] Specific examples

[1826] Example 1: Reviewing the conversation and emotions after a business meeting

[1827] The device collects and analyzes voice, image data, and emotional data during negotiations in real time. After the negotiation is over, when the user searches for the keyword "negotiation" in the application, the server finds and displays the relevant analysis data. The user can easily review details of the negotiation, information about the other party, and even their own and the other party's emotional states.

[1828] Example 2: When you want to remember the name or feelings of someone you met at an event

[1829] During the event, the device collects and analyzes audio, image data, and emotional data. After the event is over, when the user searches for "event," the server finds relevant data and displays photos, names, conversations, and emotional states. This allows the user to easily recall the names, conversations, and emotional states of people they met at the event.

[1830] An example of a prompt to actually enter

[1831] 1. "I want to review the conversation and emotional state of this business meeting."

[1832] 2. "Tell me how you would use this system to remember information about people you met at events."

[1833] This allows the system to effectively collect, analyze, store, search, and display dialogue content and related information for users, improving the user experience.

[1834] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1835] Step 1:

[1836] Language data collection

[1837] The device uses a high-performance microphone to capture surrounding audio. The input is ambient audio data, which the device collects in real time and temporarily stores in its internal memory. Then, noise filtering technology is used to remove background noise and unwanted sounds from the collected audio data. The output is clear audio data.

[1838] Step 2:

[1839] Image data collection

[1840] The device uses a camera to capture video within its field of view in real time. The input is video data of the surroundings, which the device temporarily stores in its internal memory. The image data is captured at high resolution and processed using a function to identify human faces. The output is the image data after face identification.

[1841] Step 3:

[1842] Emotion Recognition from Speech Data

[1843] The device sends the collected voice data to the emotion recognition engine. The input is clear voice data, which the emotion recognition engine analyzes and recognizes the user's emotional state by analyzing the voice tone, pitch, and rhythm. The output is the emotion analysis result.

[1844] Step 4:

[1845] Emotion recognition from image data

[1846] The device sends the collected image data to the emotion recognition engine. The input is image data after face identification, which the emotion recognition engine analyzes to recognize the user's emotional state by analyzing facial expressions and micro-movements. The output is the image emotion analysis results.

[1847] Step 5:

[1848] Language Data Analysis

[1849] The device uses voice recognition software to analyze the collected voice data. The input is clear voice data, which the voice recognition software converts into text. The software then extracts important dialogue and keywords from the text data. The output is the analyzed text data and keywords.

[1850] Step 6:

[1851] Image data analysis

[1852] Image recognition software is used to analyze the image data collected by the device. The input is image data after face identification, which the image recognition software analyzes to detect human faces and identify people by comparing them with information in an existing database. The output is the person recognition result.

[1853] Step 7:

[1854] Sending data

[1855] The terminal packages the analyzed language data, image data, and emotion data and transmits them in encrypted form to the server. The input is the analyzed language data, image data, and emotion data, which the terminal encrypts and transmits, and the output is the encrypted data transmitted to the server.

[1856] Step 8:

[1857] Data storage

[1858] The server stores the received data in a database for each user. The input is encrypted data sent to the server, which decrypts it and stores it in the database. Here, metadata such as date and time, geographic information, emotion data, and keywords are added, and an index is generated. The output is an indexed database entry.

[1859] Step 9:

[1860] Search by keyword

[1861] A user enters search keywords through a dedicated application or web portal. The input is the user-specified search keywords, which are sent to the server. The output is a search query based on the keywords.

[1862] Step 10:

[1863] Searching for Data

[1864] The server searches for relevant data from the user's database based on keywords. The input is a search query, and the server filters the relevant data based on this using date, time, geographic information, sentiment data, and indexes. The output is a dataset of search results.

[1865] Step 11:

[1866] Displaying search results

[1867] The server sends the search results to the user's device. The input is a dataset of search results, which is sent to the user's device. The user's device displays the search results based on the received data, allowing the user to visually confirm related conversation content, information about the other party, and their emotional state. The output is the displayed search results.

[1868] (Application example 2)

[1869] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1870] In traditional brick-and-mortar stores, it was difficult for store associates to understand customer interactions and their emotional state in real time, making it difficult to optimize the quality of service. Furthermore, there was a lack of effective means to make optimal suggestions and follow-ups based on customer emotions, which could result in lower customer satisfaction. Furthermore, traditional systems lacked the ability to provide real-time feedback on collected data, making it difficult for store associates to respond immediately.

[1871] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting language data, means for collecting image data, means for extracting dialogue content from the collected language data, means for identifying people from the collected image data, means for recognizing emotions from the collected language data and image data, means for transmitting the analyzed language data, image data, and emotion data to the server, means for storing the transmitted data in a database for each user, means for providing feedback on the stored data in real time, means for searching the stored data based on keywords, and means for displaying search results on the user's terminal. This allows store clerks to grasp the dialogue and emotional state of customers in real time and respond appropriately, improving the quality of service and increasing customer satisfaction.

[1872] The "means for collecting language data" refers to a device or method that uses a microphone or the like to capture surrounding sounds and collect them as language data.

[1873] "Means for collecting image data" refers to a device or method that uses a camera or the like to capture images within the field of view and collects them as image data.

[1874] The "means for extracting dialogue content from collected language data" refers to a device or method that converts collected language data into text using speech recognition technology and extracts important dialogue content and keywords.

[1875] "Means for identifying a person from collected image data" refers to a device or method that uses image recognition technology to identify a person from collected image data and compare it with a database.

[1876] The "means for recognizing emotions from collected language data and image data" refers to a device or method for recognizing the emotional state of a user by analyzing voice tone, facial expressions, etc.

[1877] The "means for transmitting analyzed language data, image data, and emotion data to a server" refers to a device or method for packaging the analyzed data and transmitting it to a server via a network.

[1878] The "means for storing transmitted data in a database for each user" refers to a device or method for sorting and storing the analysis data received by the server for each user.

[1879] "Means for providing real-time feedback of stored data" refers to a device or method that immediately processes the analyzed and stored data and displays it in real-time through a user interface.

[1880] The term "means for searching stored data based on keywords" refers to a device or method for efficiently searching related information from a stored database based on keywords entered by a user.

[1881] "Means for displaying search results on the user's terminal" refers to a device or method for transmitting searched data to the user's terminal and visually displaying it through an interface.

[1882] This invention is a system that enables store clerks in brick-and-mortar stores to grasp the conversations and emotional state of customers in real time and optimize the quality of service. The system consists of smart glasses, a voice recognition engine, an image recognition engine, an emotion recognition engine, a server, cloud storage, and a real-time display. This system allows the details of conversations with customers and the emotional state of customers to be grasped instantly, enabling the provision of appropriate services.

[1883] Hardware and Software Configuration

[1884] Smart Glasses

[1885] The smart glasses are equipped with high-performance microphones and cameras that collect voice and image data from customer interactions in real time.

[1886] Speech Recognition Engine

[1887] The speech recognition engine analyzes the collected voice data and converts it into text. This analysis is performed using, for example, the Google Speech Recognition API.

[1888] Image Recognition Engine

[1889] The image recognition engine analyzes the collected image data and identifies the customer's face using tools such as OpenCV and DeepFace.

[1890] Emotion Recognition Engine

[1891] The emotion recognition engine analyzes collected voice and image data to recognize the customer's emotional state, using the Emotion API and machine learning models for emotion recognition.

[1892] Server and Cloud Storage

[1893] The collected data and analysis results are sent in encrypted form to a server and stored in a per-user database using cloud storage services such as AWS or Google Cloud Storage.

[1894] Real-time Display

[1895] Store associates receive real-time feedback through the smart glasses' display, allowing them to respond immediately to the customer's immediate needs and emotional state.

[1896] Processing flow explanation

[1897] 1. Audio and Image Collection:

[1898] The smart glasses collect customer interaction through a microphone and capture images of the surroundings through a camera, and these data are temporarily stored in the internal memory.

[1899] 2. Data Analysis:

[1900] The voice data is converted into text by a voice recognition engine, which then analyzes emotions from the tone and pitch of the voice, while the image data is analyzed by an image recognition engine to identify the customer's face and recognize their emotional state from their facial expressions.

[1901] 3. Data transmission and storage:

[1902] The analyzed data is encrypted and sent to a server, which then stores the data in cloud storage and organizes it for each user.

[1903] 4. Real-time feedback:

[1904] The smart glasses display allows store staff to see the customer's conversation and emotional state in real time, allowing them to provide optimal responses and suggestions based on this feedback.

[1905] Specific examples

[1906] 1. Know when a customer is interested in a specific product:

[1907] The program recognizes the customer's interests in real time and provides feedback to the store clerk, saying, "You are interested in this product."

[1908] 2. Recognize when a customer is unhappy:

[1909] The program detects a customer's dissatisfied facial expression and displays a warning to the store clerk saying, "The customer is dissatisfied."

[1910] Prompt Sentence Examples

[1911] An example of a prompt is: "When a customer asks a question about a particular product, the smart glasses analyze the conversation and the customer's facial expressions in real time and display feedback such as, 'It seems you are interested in this product. Further explanation is needed.'"

[1912] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1913] Step 1:

[1914] The device (smart glasses) collects the conversational voice with the customer using a microphone and captures the surrounding field of view using a camera. These data are temporarily stored in the internal memory. Voice data and image data are collected as input and saved in the internal memory as output. In concrete terms, the microphone captures the voice signal and the camera continuously takes images.

[1915] Step 2:

[1916] The device sends the collected voice data to a speech recognition engine, which converts it into text data. The input is voice data, which the speech recognition engine converts into text. Text data is generated as output. Specifically, the device uses the Google Speech Recognition API to convert the voice signal into text.

[1917] Step 3:

[1918] The device sends the collected image data to an image recognition engine to identify the customer's face. The input is image data, which the image recognition engine analyzes to identify the person. The output is information about the identified person. Specifically, the device uses OpenCV and DeepFace to analyze facial features and compare them with a database.

[1919] Step 4:

[1920] The device sends the collected voice and image data to the emotion recognition engine to analyze the customer's emotional state. The input is voice and image data, which the emotion recognition engine analyzes to obtain emotional data. Emotion data is generated as output. Specifically, the device uses the Emotion API to analyze voice tone and facial expressions.

[1921] Step 5:

[1922] The device encrypts the analyzed language data, image data, and emotion data and sends it to the server. The input is the analyzed data, which is encrypted and sent to the server via the network. The output is the data sent to the server. Specifically, the device sends the data using an encryption protocol such as SSL / TLS.

[1923] Step 6:

[1924] The server stores the analysis data it receives in a database for each user. The input is the analyzed data, which is organized for each user and stored in the database. The output is the analysis data stored in the database. Specifically, the server organizes and stores the data using AWS or Google Cloud Storage.

[1925] Step 7:

[1926] The server sends analysis data for feedback to the device in real time. The input is analysis data obtained from the user's database, which is sent to the device in real time. The output is feedback that is displayed immediately on the device. Specifically, the server sends data immediately using a protocol such as WebSocket.

[1927] Step 8:

[1928] Based on the feedback displayed on the terminal in real time, the user (store clerk) responds appropriately to the customer. The input is feedback data from the terminal, and the user provides services based on this. The output is the result of providing an appropriate service. In concrete terms, the user checks the display on the smart glasses and engages in a dialogue based on the feedback content.

[1929] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1930] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1931] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1932] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1933] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1934] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1935] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1936] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1937] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1938] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1939] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1940] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1941] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1942] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1943] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1944] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1945] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1946] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1947] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes withou...

Claims

1. a means of collecting language data; means for collecting image data; A means for analyzing the collected language data and extracting the dialogue content; means for analyzing the collected image data to identify a person; means for transmitting the analyzed language data and image data to a server; a means for storing the transmitted data in a per-user database; a means for searching the stored data based on keywords; means for displaying search results on a user's device; A system including:

2. means for automatically generating keywords based on the analyzed linguistic data; A means of indexing the database based on automatically generated keywords; The system of claim 1 further comprising:

3. A means for comparing the analyzed image data with an existing person database; a means for automatically updating new person information; The system of claim 1 further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A