system
The system addresses inefficiencies in paper-based home visit surveys by converting audio to text, summarizing, and automatically inputting data into forms, while detecting and correcting errors, thereby improving accuracy and reliability.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
The process of recording home visit survey data on paper and pen is labor-intensive, prone to errors, and requires additional time for corrections, leading to inefficiencies and reduced accuracy.
A system that converts audio data from on-site surveys into text in real time, summarizes the information, and automatically inputs it into a survey form, with detection mechanisms to identify omissions or inconsistencies and generate additional questions.
This system enhances the efficiency and accuracy of surveys by reducing staff burden and ensuring highly reliable survey forms.
Smart Images

Figure 2026069077000001_ABST
Abstract
Description
Technical Field
[0005]
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the home visit survey for long-term care certification, data is usually recorded on paper and pen and then a survey form is created. However, this process may involve a lot of labor and errors. Also, if an omission is found during the survey, it needs to be corrected later, which takes extra time and effort. As a result, the burden on staff increases, and the efficiency and accuracy of the survey may decrease.
Means for Solving the Problems
[0005] This invention provides a system that converts audio data acquired during on-site surveys into text information in real time, summarizes that information, and automatically inputs it into a survey form. This system includes acquisition means, conversion means, summarization means, input means, detection means, and generation means, and improves the efficiency and accuracy of surveys by immediately detecting any omissions or inconsistencies in information during the survey and generating additional questions. This reduces the burden on staff and enables the creation of highly reliable survey forms.
[0006] "Acquisition method" refers to the function of recording audio during on-site surveys and acquiring it as digital data.
[0007] The "conversion means" refers to a speech recognition function that converts acquired audio data into text data.
[0008] A "summarization tool" is a function that efficiently organizes information from converted text data and extracts only the necessary content.
[0009] "Input method" refers to a function that automatically fills in summarized text information into the prescribed format of the questionnaire.
[0010] A "detection method" is a function that checks for omissions or logical inconsistencies in the input information and identifies any overlooked details.
[0011] A "generation mechanism" is a function that generates appropriate additional questions based on the deficiencies or inconsistencies in the detected information. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4]It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface that includes a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This invention relates to the construction of a system that enables the acquisition of audio data during on-site surveys and its automatic input into questionnaires. The core of the system consists of acquisition means, conversion means, summarization means, input means, detection means, and generation means. Embodiments of each means are described below.
[0034] First, once the home visit begins, the device starts acquiring audio data. At this time, the device is equipped with a high-sensitivity microphone and noise cancellation function to eliminate external noise while clearly recording the conversation between the care recipient and the staff member.
[0035] Next, the recorded audio data is converted into text in real time by a conversion mechanism within the device. The device uses an advanced speech recognition algorithm that can handle differences in dialect and intonation.
[0036] Once the text data is generated, the server activates a summarization mechanism. This mechanism extracts important information from the vast amount of text, selecting and shortening the necessary data. This creates the summarized data ready for input into the questionnaire.
[0037] The server then automatically inputs the summarized information into each field of the questionnaire. The input method uses data formatted in a standard format, arranging the information according to a template.
[0038] Furthermore, to ensure that all entered information is accurate, the server uses detection mechanisms to detect any missing or inconsistent information. This detection mechanism is achieved by utilizing AI models and comparing the data with existing databases.
[0039] Finally, if the detection means finds a problem, the server uses the generation means to form additional questions. These questions are sent to the terminal in real time to prompt the user for further confirmation.
[0040] For example, suppose during a home visit, a care recipient is asked, "What kind of exercise do you do every day?" and the answer is, "I take walks three times a week." The server summarizes this information and enters "Exercise: Walks three times a week" into the questionnaire. However, if there are any missing details regarding specific dates or other activities related to the "every day" question, the server generates additional questions to resolve the inconsistencies.
[0041] This system makes on-site surveys more efficient, allowing staff to improve the accuracy and reliability of survey results.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] The terminal activates its voice acquisition module at the start of the on-site survey and records the conversation in real time. The acquired voice data is temporarily stored on the local device.
[0045] Step 2:
[0046] After accumulating a certain amount of continuously acquired audio data in a buffer, the terminal activates speech recognition software to convert the audio into text. Noise reduction processing is applied during this conversion process to minimize misrecognition.
[0047] Step 3:
[0048] Once speech recognition is complete, the converted text data is sent to the server. The server inputs the received data into a summarization algorithm. Here, important information is extracted and a summarized text is generated.
[0049] Step 4:
[0050] Upon receiving the summarized text information, the server begins the automatic input process into the questionnaire template. During this process, the questionnaire is constructed sequentially, mapping each summarized piece of information to its corresponding survey item.
[0051] Step 5:
[0052] To verify the entered survey data, the server uses detection methods to compare it with existing data and templates, identifying any missing or inconsistent data. Anomalies are detected using AI models during this process.
[0053] Step 6:
[0054] Based on the identified omissions and inconsistencies, the server generates additional questions using a generation mechanism. The generated questions are sent to the terminal in a timely manner and provided as real-time feedback to the user.
[0055] Step 7:
[0056] Based on additional questions presented on the device, the user reconfirms information with the care recipient and records their responses. The newly acquired audio data is then recognized and processed again and reflected in the questionnaire.
[0057] Step 8:
[0058] Finally, the server fully verifies the contents of the questionnaire, and if there are no problems, it saves the completed questionnaire to secure storage. As a result, records of on-site surveys can be managed efficiently and accurately.
[0059] (Example 1)
[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0061] In field surveys, it is necessary to improve the accuracy and efficiency of acquiring audio data accurately and efficiently, and then automatically inputting that data into questionnaires. Furthermore, a challenge is to enhance the reliability of survey results by reliably detecting missing or inconsistent information and quickly acquiring additional information.
[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0063] In this invention, the server includes means for acquiring audio data during a visit, means for converting the acquired audio data into text information, and means for summarizing the converted text information and extracting important information. This enables efficient acquisition of audio data during on-site surveys and automatic input of accurate information based on that data.
[0064] A "device for acquiring audio data during a visit" refers to a terminal equipped with a high-sensitivity microphone and noise cancellation function to effectively eliminate external noise and record conversations clearly.
[0065] A "device that converts acquired audio data into text information" refers to a terminal equipped with a speech recognition algorithm that utilizes a generative AI model to convert audio data into accurate text data in real time.
[0066] A "device for summarizing converted text information and extracting important information" refers to a server that automatically extracts important information from generated text data and shortens it to suit the purpose of the investigation.
[0067] A "device that automatically inputs data into a standardized format" refers to a server equipped with data entry capabilities to neatly arrange summarized information onto a questionnaire based on a template.
[0068] A "device for detecting omissions and inconsistencies" refers to a server equipped with an AI model that identifies problems by comparing automatically entered data with existing databases to verify its integrity.
[0069] A "device for generating questions to obtain additional necessary information" refers to a server that, based on detected omissions and inconsistencies, generates new questions in real time as needed and sends them to the terminal.
[0070] This invention provides a system for streamlining the automatic acquisition of voice data and automatic input into questionnaires during on-site surveys. A specific embodiment of this system is described below.
[0071] Once the home visit begins, the device uses a high-sensitivity microphone and noise cancellation function to clearly record conversations between the care recipient and staff. This makes it possible to obtain accurate audio data while eliminating external noise.
[0072] Next, to convert the acquired audio data into text data in real time, the device uses an advanced speech recognition algorithm that applies a generative AI model. This algorithm takes into account dialects and the speaker's intonation, enabling it to generate accurate text information.
[0073] The generated text data is sent to a server where it is summarized. The server uses natural language processing techniques to extract important information from the large amount of text, narrowing down the necessary data. This removes redundant information and ensures that the summarized data necessary for the investigation is secured.
[0074] Next, the server automatically enters the summarized information into each item of the questionnaire based on a standardized format. This input process is designed to ensure that the data is precisely placed in the correct location, increasing the efficiency of automated entry.
[0075] To detect missing or inconsistent data, the server utilizes detection mechanisms. These mechanisms compare the data with existing databases and use AI models to verify its accuracy. This improves the reliability of the investigation results.
[0076] Finally, if any omissions or inconsistencies are found in the information, the server uses a generation mechanism to formulate additional questions and send them to the terminal. This allows the user to immediately verify and provide additional information.
[0077] For example, suppose during a home visit, a care recipient is asked, "What kind of exercise do you do every day?" and the answer is, "I take walks three times a week." The server summarizes this information as "Exercise: Walks three times a week" and enters it into the questionnaire. As a result, the efficiency and accuracy of the survey are improved.
[0078] An example of a prompt to input into a generative AI model is, "Please tell me how to convert audio data obtained from a field survey into text, summarize the important information, and input it into a survey form."
[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0080] Step 1:
[0081] The terminal begins acquiring audio data simultaneously with the start of the home visit. A high-sensitivity microphone captures external conversational audio as input. By utilizing noise cancellation technology to remove unwanted background noise, clear conversation data between the care recipient and staff is output. This audio data is a crucial source of information that forms the basis of processing.
[0082] Step 2:
[0083] The device activates its built-in speech recognition software to convert the acquired audio data into text. Audio data is used as input, and advanced algorithms from a generative AI model are applied to handle speech characteristics such as dialects and intonation. The output of this process is accurate, structured text data.
[0084] Step 3:
[0085] The server receives text data and uses summarization techniques. Upon receiving text data, it utilizes natural language processing techniques to extract important information from the vast amount of data. At this stage, it removes extraneous data and outputs a shortened summary necessary for the investigation.
[0086] Step 4:
[0087] The server automatically inputs the summarized information according to the standard questionnaire format. In this process, the questionnaire template is used as input, and the summarized information is organized so that it corresponds to each item. The output is properly completed questionnaire data.
[0088] Step 5:
[0089] The server activates detection mechanisms to verify the integrity of the input data. The input consists of completed questionnaires, and by combining an AI model and a database, it identifies omissions and inconsistencies. If problems are detected during this process, the findings are provided as output.
[0090] Step 6:
[0091] When a problem is detected, the server uses a generation mechanism to form any necessary additional questions. The output from the detection mechanism becomes the input, and the generated questions are output. These questions are sent to the terminal in real time, requesting additional information from the user and ensuring data integrity with the newly obtained information.
[0092] (Application Example 1)
[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0094] Traditional in-person surveys faced challenges in acquiring voice data and efficiently inputting that data, as well as difficulty in understanding customer needs in real time and effectively recommending products. This hindered improvements in survey accuracy and customer experience.
[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0096] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting audio into text information, and a recommendation means for acquiring conversations with customers and making product recommendations. This enables improved accuracy of on-site surveys and an enhanced customer experience.
[0097] "Acquisition means" refers to a device or system for accurately acquiring audio data during a visit, while taking ambient noise into consideration.
[0098] "Conversion means" refers to the process or device for converting acquired audio data into a different format, specifically textual information.
[0099] A "summarization tool" is a process or system for extracting necessary information from converted textual information and displaying it in a shortened form.
[0100] An "input method" is a system for automatically inputting summarized information according to a standardized format.
[0101] A "detection means" is a system that has the function of comparing input data with a template to discover missing or inconsistent information.
[0102] A "recommendation method" is a system that selects and presents the most suitable products based on information obtained from conversations with customers.
[0103] A "generation means" is a process or system that has the function of generating additional questions based on the detected problem.
[0104] This invention relates to the implementation of a system for inputting information from voice data obtained through on-site surveys and in physical stores, as well as a product recommendation system. Efficient and highly accurate information processing is achieved through the cooperation of a server and terminals.
[0105] The device records conversations with customers during visits or within stores using a high-sensitivity microphone and obtains clear audio data by employing technology to remove background noise. The recorded audio is converted into text in real time using speech recognition software (e.g., Google® Speech-to-Text API) on the device. This process supports a wide range of speech, including intonation and dialects.
[0106] Subsequently, the server receives the converted text information and summarizes it using a natural language processing engine (e.g., spaCy). It extracts key information and condenses it into the information that forms the basis for each item on the questionnaire and for product recommendations. The summarized information is then formatted into a standard format and automatically entered into the questionnaire and product database.
[0107] After input, the server activates a detection mechanism and uses an AI model (e.g., TENSORFLOW®) to detect missing or inconsistent data. During this process, it compares the data with an existing database, generates additional questions as needed, and sends them to the terminal. The user can then use the terminal to respond to these additional questions for further verification.
[0108] For example, in a physical store, if a customer asks, "I'm having a barbecue with friends next Sunday, what do I need?", the server analyzes the conversation and recommends products such as "barbecue meat" and "seasoning sets," and also provides relevant discount information.
[0109] Examples of prompts for a generative AI model:
[0110] "The customer is looking for barbecue-related products. Please generate a list of required items and a list of recommended products."
[0111] In this way, it is possible to improve the efficiency and accuracy of on-site surveys, as well as enhance the customer experience at physical stores.
[0112] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0113] Step 1:
[0114] The device uses a high-sensitivity microphone to capture ambient sound and noise cancellation to obtain clear audio data. It receives conversations between customers and staff as input and outputs a clear audio file, which is then used for subsequent processing.
[0115] Step 2:
[0116] The device converts acquired audio data into text in real time using speech recognition software (e.g., Google Speech-to-Text API). It analyzes the input audio data and outputs the corresponding text data. This process also takes into account differences in intonation and dialect.
[0117] Step 3:
[0118] The server receives text data sent from the terminal and uses a natural language processing engine (e.g., spaCy) to perform summarization. It extracts important information and removes unnecessary parts to output summarized text. This summarized data is then used in the next step.
[0119] Step 4:
[0120] The server automatically formats the summarized text information into a standard format and inputs it into the questionnaire or product database. The input summarized information is then formatted, and formatted data is obtained as output. This data is then adapted to a known template.
[0121] Step 5:
[0122] The server uses an AI model (e.g., TensorFlow) to detect missing or inconsistent data in the input. This AI model compares the input data against an existing database to verify its integrity. If problems are found, the detection results are provided as output.
[0123] Step 6:
[0124] The server generates additional questions based on the detected issues and sends them to the terminal. This generation AI model is used to create specific question prompts. The additional questions are then provided to the user as output.
[0125] Step 7:
[0126] The user follows the additional questions displayed on the terminal and enters their responses. The terminal receives the entered responses and sends the data to the server. This prepares the server to supply highly accurate data as output for the next processing step.
[0127] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0128] This invention relates to a system that not only acquires voice data during on-site surveys and automatically inputs it into questionnaires, but also incorporates an emotion engine that recognizes the user's emotional state. This system includes acquisition means, conversion means, summarization means, input means, detection means, generation means, and an emotion engine. Detailed embodiments of each means and the emotion engine are described below.
[0129] To begin, the device activates its voice acquisition module at the start of the visit and records the conversation in real time. At this stage, noise cancellation technology is used to minimize ambient noise while acquiring clear voice data.
[0130] The recorded audio is immediately converted into text information by the device's built-in conversion mechanism. This conversion process utilizes diverse language models to accommodate dialects and intonation differences, achieving highly accurate text conversion.
[0131] Next, the server receives the converted text and uses a summarization tool to extract and summarize the important information. The generated summary information is automatically entered into the questionnaire based on a standardized format.
[0132] A key feature of this system is its emotion engine, which simultaneously analyzes voice data and evaluates the emotional state of both the user and the person being cared for. The emotion engine analyzes voice tone, speaking style, and context to grasp changes in emotional state in real time.
[0133] The server uses detection methods to check for any missing or inconsistent data in the entered questionnaire. During this process, it also utilizes emotional information obtained by the emotion engine to adjust the questions based on whether the user is experiencing stress or based on their responses.
[0134] Furthermore, if the detection means identifies any omissions or inconsistencies in the input data, the server uses the generation means to generate additional questions and notifies the user via the terminal.
[0135] For example, if a person receiving care gives a vague answer to the question, "Do you enjoy the food?", the server will consider their emotional state and suggest follow-up questions such as, "What kind of food do you particularly like?"
[0136] This system not only improves the accuracy of surveys but also reduces the psychological burden on users, supporting the smooth and reliable conduct of surveys.
[0137] The following describes the processing flow.
[0138] Step 1:
[0139] The device activates its voice acquisition function as soon as the on-site survey begins, recording the conversation clearly. Recording takes place in the background, and ambient noise is removed using the noise cancellation function.
[0140] Step 2:
[0141] The device transfers the acquired audio data to the speech recognition system in real time, where it is converted into text. The converted text is then saved in a format that can be processed immediately.
[0142] Step 3:
[0143] The server receives text data sent from the terminal and applies a summarization algorithm. This algorithm efficiently extracts important information and generates a summary.
[0144] Step 4:
[0145] The server automatically inputs the generated summary information into the questionnaire according to a standard format. Here, the information is accurately mapped to a pre-configured template.
[0146] Step 5:
[0147] Simultaneously, the device's emotion engine analyzes the conversation between the user and the person being cared for, recognizing their emotional state from their tone of voice and the flow of the conversation. Emotional data is accumulated in real time.
[0148] Step 6:
[0149] The server verifies the submitted questionnaires to check for any omissions or inconsistencies. During this process, it utilizes data from the emotion engine, focusing particularly on areas where emotional responses were strong.
[0150] Step 7:
[0151] If any omissions or inconsistencies are detected, the server will consider the emotional context and generate appropriate additional questions. The generated questions will be notified to the user via the terminal.
[0152] Step 8:
[0153] The user follows the instructions on the device, asks additional questions to the person being cared for, and records their responses. The newly obtained audio data is processed according to the aforementioned flow and reflected in the questionnaire.
[0154] Step 9:
[0155] Finally, the server verifies the content and securely stores the complete questionnaire in cloud storage or a database. This improves the accuracy of the survey records.
[0156] (Example 2)
[0157] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0158] The aim is to streamline the process of home visit surveys, from acquiring audio data to inputting information and analyzing emotional states, thereby ensuring accuracy and smooth progress of the surveys. Furthermore, the goal is to provide a system that reduces the psychological burden on care recipients and surveyors during the survey, enabling the acquisition of highly reliable data.
[0159] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0160] In this invention, the server includes an acquisition means for acquiring audio during a visit, a conversion means for converting audio into text information, a summarization means for extracting and summarizing important information, an emotion analysis means for evaluating emotional states through audio data analysis, and a generation means for verifying input data and generating additional inquiries. This enables efficient and accurate data acquisition during on-site surveys, as well as flexible survey progress that takes emotional states into consideration.
[0161] "Acquisition method" refers to a function for collecting audio data during visits while suppressing environmental noise.
[0162] The "conversion means" is a function that converts acquired speech into text information, and it uses a highly accurate language model.
[0163] A "summarization tool" is a function that extracts important information from text and converts it into a summarized format.
[0164] "Input method" refers to a function that automatically inputs summarized information into a standardized information format.
[0165] "Emotion analysis means" refers to a function that evaluates the emotional state of the user and the person they are talking to based on voice data.
[0166] A "detection mechanism" is a function that verifies input data, checks for omissions or inconsistencies, and adjusts the information.
[0167] A "generation mechanism" is a function that generates additional inquiries based on detected problems and missing information, and presents them to the interlocutor.
[0168] To implement this invention, a system is constructed that streamlines the on-site survey process and improves its accuracy by using the following hardware and software.
[0169] First, the device is equipped with an audio acquisition module and uses noise cancellation technology to acquire clear audio data. High-performance audio capture devices and noise reduction software are used for audio acquisition. For example, a microphone built into a smart device carried by the researcher and a signal processing algorithm that suppresses ambient noise could be considered.
[0170] Next, the device converts the audio data into text information using a conversion method. A cloud-based speech recognition API is used for speech recognition, and a multilingual language model handles dialects and intonation differences. For example, the Google Cloud Speech-to-Text API is used to convert audio data into text.
[0171] The server analyzes the received text information using a summarization mechanism, extracting and summarizing important information. Using natural language processing techniques, a specific algorithm automatically summarizes the key points of the information. For example, an algorithm that analyzes word frequency and relationships is used for information extraction.
[0172] On the other hand, the emotion analysis system analyzes voice data and evaluates the emotions of the user and the person they are speaking with. An emotion recognition algorithm is implemented to analyze the emotional state in real time based on the tone and content of the voice. This allows for the evaluation of changes in the emotions of the person being surveyed during the investigation.
[0173] The server uses detection methods to verify the input data and compare it to a template to check for any omissions or inconsistencies. By also considering the results of sentiment analysis, it is possible to adjust the questions if the user is experiencing stress.
[0174] Finally, the server generates additional queries to complement the detected problems using a generation mechanism. Using a generation AI model, it creates new questions based on appropriate prompts and presents them to the user via the terminal. An example of a prompt might be, "In a home visit, please describe how to analyze elderly people's feelings about food and identify their specific preferences."
[0175] This system improves the accuracy of on-site surveys, reduces the psychological burden on both the user and the interviewee, and enables smooth and reliable surveys.
[0176] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0177] Step 1:
[0178] The device activates its voice acquisition module at the start of the on-site survey and records voice data in real time while suppressing ambient noise using noise cancellation technology. The input is the conversation with the survey subject, and the output is a clear audio file. Specifically, the device's recording function is activated, and noise filtering is performed in the background.
[0179] Step 2:
[0180] The device converts recorded audio into text information using a conversion mechanism. The input is an audio file, and the output is text. Using a speech recognition API, a language model analyzes the audio waveform to convert it into text. Specifically, the audio is sent to a cloud service, and each audio segment is mapped to its corresponding text.
[0181] Step 3:
[0182] The server receives the converted text information, extracts important information using a summarization tool, and performs a summary. The input is text, and the output is summarized information. Natural language processing technology is used to analyze keywords and themes from the text and perform a summary. Specifically, a text analysis algorithm scans the content and selects sentences of high importance.
[0183] Step 4:
[0184] The server automatically inputs summarized information into a standardized information format. The input is summarized information, and the output is a standardized survey record. A database management system is used, and the summarization results are stored in the database according to the fields. Specifically, the information is automatically entered according to a predetermined format and registered as a survey record.
[0185] Step 5:
[0186] The emotion analysis tool analyzes the audio data and evaluates the emotional state of the user and the person they are speaking with. The input is the audio data, and the output is a score indicating the emotional state. The emotion identification algorithm analyzes the tone and pitch of the voice to identify changes in emotion. Specifically, the audio feature extraction tool scans the audio and assigns emotion labels.
[0187] Step 6:
[0188] The server uses detection mechanisms to check for missing or inconsistent data in the input and adjusts the data to take sentiment into account. The input consists of standardized survey records and sentiment scores, and the output is the corrected survey records. If inconsistencies are detected, the system re-evaluates the data and generates adjustment proposals. Specifically, the algorithm identifies the locations of missing data and applies corrective processing as needed.
[0189] Step 7:
[0190] The server uses a generation mechanism to generate additional inquiries regarding detected problems and notifies the user via the terminal. The input is inconsistent data and sentiment information, and the output is the additional inquiries. A generation AI model creates prompts and generates the necessary additional questions. For example, if there is an ambiguous answer, a question such as "What specific dishes do you like?" is automatically generated and suggested to the user via the terminal.
[0191] (Application Example 2)
[0192] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0193] In many physical stores, communication between customers and service staff significantly impacts the quality of service. However, it is difficult for service staff to accurately grasp the emotions and needs of all customers in real time, leading to inconsistencies in service quality. Furthermore, there are insufficient resources available to improve customer satisfaction by appropriately analyzing and addressing customer emotions and needs.
[0194] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0195] In this invention, the server includes a collection means for acquiring audio during a visit, a conversion means for converting the collected audio into text information, and an emotion analysis means for analyzing the audio and evaluating the emotional state. This makes it possible to improve the quality of communication between customers and service staff in physical stores, understand the customer's emotional state, and improve the response.
[0196] "Collection means" refers to a mechanism for acquiring voice data during visits.
[0197] "Conversion means" refers to technology for converting acquired audio data into text information.
[0198] A "summarization method" is a technique that extracts and summarizes important information from converted textual information.
[0199] "Input method" refers to technology that automatically inputs summarized information into a standard format.
[0200] A "detection means" is a mechanism for detecting omissions or inconsistencies in the input data.
[0201] "Generation means" refers to techniques for generating additional questions based on detected problems.
[0202] "Emotional analysis methods" refer to technologies that analyze voice to evaluate emotional states.
[0203] "Suggestion methods" refer to methods for improving customer service based on information obtained through emotion analysis.
[0204] The embodiment for carrying out the invention is configured as follows as a method for realizing a system to efficiently improve customer service in physical stores. This system has the function of collecting voice in real time, converting it into text information, and summarizing it. In addition, it analyzes the emotional state of customers from the voice data and uses that information to make suggestions for improving customer service.
[0205] The server uses speech recognition software to collect conversations between customers and service staff during visits. This collection method utilizes a recording device employing noise-canceling technology. Next, the collected audio data is converted into text information using natural language processing. At this stage, dialects and intonation are taken into consideration to generate accurate text data.
[0206] The converted textual information is summarized by a summarization algorithm, which extracts the most important information. This summarized information is then automatically entered according to a predetermined format. After input, the server uses detection means to check for any missing or inconsistent data, and generates additional questions as needed using generation means.
[0207] Furthermore, the server utilizes emotion analysis techniques to analyze text derived from the audio data and assess the customer's emotional state. This analysis employs a generative AI model that captures emotional changes based on voice tone and context.
[0208] For example, if a customer expresses positive feelings towards a product, the server can use this information to make further suggestions to the customer service representative, offering more helpful service. An example of a prompt would be, "Please provide sentiment analysis and summary of this conversation: 'I liked this product, but I have one more question.'"
[0209] This will improve the quality of customer service at physical stores and increase customer satisfaction.
[0210] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0211] Step 1:
[0212] The terminal activates the voice acquisition module and collects conversations between customers and service staff in real time during visits. Using noise cancellation technology, it acquires clear voice data while minimizing ambient noise. An audio signal is provided as input, and noise-reduced voice data is obtained as output.
[0213] Step 2:
[0214] The server processes the collected audio data using speech recognition software and converts it into text information. Speech recognition technology is used to convert audio information into text data. At this stage, the input is denoised audio data, and the output is recognized text data.
[0215] Step 3:
[0216] The server processes the converted character information using a summarization algorithm, extracting important information and creating a summary. This process uses natural language processing techniques to extract key points and shorten the text. The input is recognized character data, and the output is summarized text information.
[0217] Step 4:
[0218] The server automatically inputs summarized information into a pre-configured format. This format has a fixed data structure, and the data is automatically entered according to it. The input is summarized text information, and the output is a pre-filled datasheet according to the format.
[0219] Step 5:
[0220] The server analyzes the input data using detection methods to check for any missing or inconsistent data. The purpose of this process is to verify data integrity and identify any problems. The input is formatted data, and the output is the result of data verification.
[0221] Step 6:
[0222] Based on the problems detected by the server, the generation mechanism generates additional questions and notifies the user via the terminal. Here, questions are dynamically created if there are problems. The input is the verification result, and the output is a set of follow-up questions.
[0223] Step 7:
[0224] The server uses emotion analysis tools to evaluate the emotional state of text obtained from speech. A generative AI model is used to analyze voice tone and context. The input is textual information, and the output is emotional state data.
[0225] Step 8:
[0226] Based on the results of sentiment analysis, the server activates a suggestion mechanism to propose service improvements to the user. The results are returned as suggestions for the service to the user. The input is sentiment state data, and the output is improvement suggestions.
[0227] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0228] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0229] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0230] [Second Embodiment]
[0231] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0232] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0233] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0234] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0235] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0236] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0237] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0238] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0239] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0240] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0241] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0242] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0243] This invention relates to the construction of a system that enables the acquisition of audio data during on-site surveys and its automatic input into questionnaires. The core of the system consists of acquisition means, conversion means, summarization means, input means, detection means, and generation means. Embodiments of each means are described below.
[0244] First, once the home visit begins, the device starts acquiring audio data. At this time, the device is equipped with a high-sensitivity microphone and noise cancellation function to eliminate external noise while clearly recording the conversation between the care recipient and the staff member.
[0245] Next, the recorded audio data is converted into text in real time by a conversion mechanism within the device. The device uses an advanced speech recognition algorithm that can handle differences in dialect and intonation.
[0246] Once the text data is generated, the server activates a summarization mechanism. This mechanism extracts important information from the vast amount of text, selecting and shortening the necessary data. This creates the summarized data ready for input into the questionnaire.
[0247] The server then automatically inputs the summarized information into each field of the questionnaire. The input method uses data formatted in a standard format, arranging the information according to a template.
[0248] Furthermore, to ensure that all entered information is accurate, the server uses detection mechanisms to detect any missing or inconsistent information. This detection mechanism is achieved by utilizing AI models and comparing the data with existing databases.
[0249] Finally, if the detection means finds a problem, the server uses the generation means to form additional questions. These questions are sent to the terminal in real time to prompt the user for further confirmation.
[0250] For example, suppose during a home visit, a care recipient is asked, "What kind of exercise do you do every day?" and the answer is, "I take walks three times a week." The server summarizes this information and enters "Exercise: Walks three times a week" into the questionnaire. However, if there are any missing details regarding specific dates or other activities related to the "every day" question, the server generates additional questions to resolve the inconsistencies.
[0251] This system makes on-site surveys more efficient, allowing staff to improve the accuracy and reliability of survey results.
[0252] The following describes the processing flow.
[0253] Step 1:
[0254] The terminal activates its voice acquisition module at the start of the on-site survey and records the conversation in real time. The acquired voice data is temporarily stored on the local device.
[0255] Step 2:
[0256] After accumulating a certain amount of continuously acquired audio data in a buffer, the terminal activates speech recognition software to convert the audio into text. Noise reduction processing is applied during this conversion process to minimize misrecognition.
[0257] Step 3:
[0258] Once speech recognition is complete, the converted text data is sent to the server. The server inputs the received data into a summarization algorithm. Here, important information is extracted and a summarized text is generated.
[0259] Step 4:
[0260] Upon receiving the summarized text information, the server begins the automatic input process into the questionnaire template. During this process, the questionnaire is constructed sequentially, mapping each summarized piece of information to its corresponding survey item.
[0261] Step 5:
[0262] To verify the entered survey data, the server uses detection methods to compare it with existing data and templates, identifying any missing or inconsistent data. Anomalies are detected using AI models during this process.
[0263] Step 6:
[0264] Based on the identified omissions and inconsistencies, the server generates additional questions using a generation mechanism. The generated questions are sent to the terminal in a timely manner and provided as real-time feedback to the user.
[0265] Step 7:
[0266] Based on additional questions presented on the device, the user reconfirms information with the care recipient and records their responses. The newly acquired audio data is then recognized and processed again and reflected in the questionnaire.
[0267] Step 8:
[0268] Finally, the server fully verifies the contents of the questionnaire, and if there are no problems, it saves the completed questionnaire to secure storage. As a result, records of on-site surveys can be managed efficiently and accurately.
[0269] (Example 1)
[0270] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0271] In field surveys, it is necessary to improve the accuracy and efficiency of acquiring audio data accurately and efficiently, and then automatically inputting that data into questionnaires. Furthermore, a challenge is to enhance the reliability of survey results by reliably detecting missing or inconsistent information and quickly acquiring additional information.
[0272] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0273] In this invention, the server includes means for acquiring audio data during a visit, means for converting the acquired audio data into text information, and means for summarizing the converted text information and extracting important information. This enables efficient acquisition of audio data during on-site surveys and automatic input of accurate information based on that data.
[0274] A "device for acquiring audio data during a visit" refers to a terminal equipped with a high-sensitivity microphone and noise cancellation function to effectively eliminate external noise and record conversations clearly.
[0275] A "device that converts acquired audio data into text information" refers to a terminal equipped with a speech recognition algorithm that utilizes a generative AI model to convert audio data into accurate text data in real time.
[0276] A "device for summarizing converted text information and extracting important information" refers to a server that automatically extracts important information from generated text data and shortens it to suit the purpose of the investigation.
[0277] A "device that automatically inputs data into a standardized format" refers to a server equipped with data entry capabilities to neatly arrange summarized information onto a questionnaire based on a template.
[0278] A "device for detecting omissions and inconsistencies" refers to a server equipped with an AI model that identifies problems by comparing automatically entered data with existing databases to verify its integrity.
[0279] A "device for generating questions to obtain additional necessary information" refers to a server that, based on detected omissions and inconsistencies, generates new questions in real time as needed and sends them to the terminal.
[0280] This invention provides a system for streamlining the automatic acquisition of voice data and automatic input into questionnaires during on-site surveys. A specific embodiment of this system is described below.
[0281] Once the home visit begins, the device uses a high-sensitivity microphone and noise cancellation function to clearly record conversations between the care recipient and staff. This makes it possible to obtain accurate audio data while eliminating external noise.
[0282] Next, to convert the acquired audio data into text data in real time, the device uses an advanced speech recognition algorithm that applies a generative AI model. This algorithm takes into account dialects and the speaker's intonation, enabling it to generate accurate text information.
[0283] The generated text data is sent to the server and summarized. The server uses natural language processing technology to extract important information from large-scale text and narrow down the necessary data. As a result, redundant information is removed, and the summary data required for the investigation is ensured.
[0284] Next, the server automatically inputs the summarized information into each item of the questionnaire based on a standardized format. This input process is designed so that the data is accurately placed in the correct position, enhancing the efficiency of automatic input.
[0285] To detect missing or inconsistent input data, the server utilizes detection means. This means compares with an existing database and uses an AI model to confirm the accuracy of the information. As a result, the reliability of the investigation results is improved.
[0286] Finally, if missing or inconsistent information is found in the information, the server uses generation means to form additional questions and send them to the terminal. As a result, the user can immediately confirm and provide additional information.
[0287] For example, suppose that during a visit, the caregiver asks "What kind of exercise do you do every day?", and the answer is the information "I take a walk three times a week". The server summarizes this information as "Exercise: Walk three times a week" and inputs it into the questionnaire. As a result, the efficiency and accuracy of the investigation are improved.
[0288] An example of a prompt sentence input into the generation AI model is "Please teach me how to convert the voice data obtained from a door-to-door survey into text, summarize the important information, and input it into a questionnaire."
[0289] The flow of specific processing in Example 1 will be described using FIG. 11.
[0290] Step 1:
[0291] The terminal begins acquiring audio data simultaneously with the start of the home visit. A high-sensitivity microphone captures external conversational audio as input. By utilizing noise cancellation technology to remove unwanted background noise, clear conversation data between the care recipient and staff is output. This audio data is a crucial source of information that forms the basis of processing.
[0292] Step 2:
[0293] The device activates its built-in speech recognition software to convert the acquired audio data into text. Audio data is used as input, and advanced algorithms from a generative AI model are applied to handle speech characteristics such as dialects and intonation. The output of this process is accurate, structured text data.
[0294] Step 3:
[0295] The server receives text data and uses summarization techniques. Upon receiving text data, it utilizes natural language processing techniques to extract important information from the vast amount of data. At this stage, it removes extraneous data and outputs a shortened summary necessary for the investigation.
[0296] Step 4:
[0297] The server automatically inputs the summarized information according to the standard questionnaire format. In this process, the questionnaire template is used as input, and the summarized information is organized so that it corresponds to each item. The output is properly completed questionnaire data.
[0298] Step 5:
[0299] The server activates detection mechanisms to verify the integrity of the input data. The input consists of completed questionnaires, and by combining an AI model and a database, it identifies omissions and inconsistencies. If problems are detected during this process, the findings are provided as output.
[0300] Step 6:
[0301] When the server detects a problem, it forms the necessary additional questions using the generation means. The output from the detection means serves as the input, and the generated questions are output. These questions are sent to the terminal in real time, asking the user for additional information and ensuring data consistency with the newly obtained information.
[0302] (Application Example 1)
[0303] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0304] In conventional door-to-door surveys, there are problems in obtaining voice data and efficient information input based on it. Furthermore, it is difficult to grasp customers' needs in real time and effectively recommend products. As a result, improving the accuracy of surveys and the customer experience has been hindered.
[0305] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0306] In this invention, the server includes an acquisition means for acquiring voice, a conversion means for converting voice into character information, and a recommendation means for acquiring conversations with customers and making product recommendations. This enables improving the accuracy of door-to-door surveys and the customer experience.
[0307] The "acquisition means" is a device or system for accurately acquiring voice data during a visit while considering environmental noise.
[0308] The "conversion means" is a process or device for converting the acquired voice data into a different format, specifically character information.
[0309] The "summarization means" is a process or system for extracting necessary information from the converted character information and displaying it in a shortened form.
[0310] An "input method" is a system for automatically inputting summarized information according to a standardized format.
[0311] A "detection means" is a system that has the function of comparing input data with a template to discover missing or inconsistent information.
[0312] A "recommendation method" is a system that selects and presents the most suitable products based on information obtained from conversations with customers.
[0313] A "generation means" is a process or system that has the function of generating additional questions based on the detected problem.
[0314] This invention relates to the implementation of a system for inputting information from voice data obtained through on-site surveys and in physical stores, as well as a product recommendation system. Efficient and highly accurate information processing is achieved through the cooperation of a server and terminals.
[0315] The device records conversations with customers during visits or within stores using a high-sensitivity microphone and obtains clear audio data by employing technology to remove background noise. The recorded audio is converted into text in real time using speech recognition software (e.g., Google Speech-to-Text API) on the device. This process also supports diverse speech patterns, including intonation and dialects.
[0316] Subsequently, the server receives the converted text information and summarizes it using a natural language processing engine (e.g., spaCy). It extracts key information and condenses it into the information that forms the basis for each item on the questionnaire and for product recommendations. The summarized information is then formatted into a standard format and automatically entered into the questionnaire and product database.
[0317] After input, the server activates a detection mechanism and uses an AI model (e.g., TensorFlow) to detect missing or inconsistent data. During this process, it compares the data with an existing database, generates additional questions as needed, and sends them to the terminal. The user can then use the terminal to answer these additional questions for further verification.
[0318] For example, in a physical store, if a customer asks, "I'm having a barbecue with friends next Sunday, what do I need?", the server analyzes the conversation and recommends products such as "barbecue meat" and "seasoning sets," and also provides relevant discount information.
[0319] Examples of prompts for a generative AI model:
[0320] "The customer is looking for barbecue-related products. Please generate a list of required items and a list of recommended products."
[0321] In this way, it is possible to improve the efficiency and accuracy of on-site surveys, as well as enhance the customer experience at physical stores.
[0322] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0323] Step 1:
[0324] The device uses a high-sensitivity microphone to capture ambient sound and noise cancellation to obtain clear audio data. It receives conversations between customers and staff as input and outputs a clear audio file, which is then used for subsequent processing.
[0325] Step 2:
[0326] The device converts acquired audio data into text in real time using speech recognition software (e.g., Google Speech-to-Text API). It analyzes the input audio data and outputs the corresponding text data. This process also takes into account differences in intonation and dialect.
[0327] Step 3:
[0328] The server receives text data sent from the terminal and uses a natural language processing engine (e.g., spaCy) to perform summarization. It extracts important information and removes unnecessary parts to output summarized text. This summarized data is then used in the next step.
[0329] Step 4:
[0330] The server automatically formats the summarized text information into a standard format and inputs it into the questionnaire or product database. The input summarized information is then formatted, and formatted data is obtained as output. This data is then adapted to a known template.
[0331] Step 5:
[0332] The server uses an AI model (e.g., TensorFlow) to detect missing or inconsistent data in the input. This AI model compares the input data against an existing database to verify its integrity. If problems are found, the detection results are provided as output.
[0333] Step 6:
[0334] The server generates additional questions based on the detected issues and sends them to the terminal. This generation AI model is used to create specific question prompts. The additional questions are then provided to the user as output.
[0335] Step 7:
[0336] The user follows the additional questions displayed on the terminal and enters their responses. The terminal receives the entered responses and sends the data to the server. This prepares the server to supply highly accurate data as output for the next processing step.
[0337] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0338] This invention relates to a system that not only acquires voice data during on-site surveys and automatically inputs it into questionnaires, but also incorporates an emotion engine that recognizes the user's emotional state. This system includes acquisition means, conversion means, summarization means, input means, detection means, generation means, and an emotion engine. Detailed embodiments of each means and the emotion engine are described below.
[0339] To begin, the device activates its voice acquisition module at the start of the visit and records the conversation in real time. At this stage, noise cancellation technology is used to minimize ambient noise while acquiring clear voice data.
[0340] The recorded audio is immediately converted into text information by the device's built-in conversion mechanism. This conversion process utilizes diverse language models to accommodate dialects and intonation differences, achieving highly accurate text conversion.
[0341] Next, the server receives the converted text and uses a summarization tool to extract and summarize the important information. The generated summary information is automatically entered into the questionnaire based on a standardized format.
[0342] A key feature of this system is its emotion engine, which simultaneously analyzes voice data and evaluates the emotional state of both the user and the person being cared for. The emotion engine analyzes voice tone, speaking style, and context to grasp changes in emotional state in real time.
[0343] The server uses detection methods to check for any missing or inconsistent data in the entered questionnaire. During this process, it also utilizes emotional information obtained by the emotion engine to adjust the questions based on whether the user is experiencing stress or based on their responses.
[0344] Furthermore, if the detection means identifies any omissions or inconsistencies in the input data, the server uses the generation means to generate additional questions and notifies the user via the terminal.
[0345] For example, if a person receiving care gives a vague answer to the question, "Do you enjoy the food?", the server will consider their emotional state and suggest follow-up questions such as, "What kind of food do you particularly like?"
[0346] This system not only improves the accuracy of surveys but also reduces the psychological burden on users, supporting the smooth and reliable conduct of surveys.
[0347] The following describes the processing flow.
[0348] Step 1:
[0349] The device activates its voice acquisition function as soon as the on-site survey begins, recording the conversation clearly. Recording takes place in the background, and ambient noise is removed using the noise cancellation function.
[0350] Step 2:
[0351] The device transfers the acquired audio data to the speech recognition system in real time, where it is converted into text. The converted text is then saved in a format that can be processed immediately.
[0352] Step 3:
[0353] The server receives text data sent from the terminal and applies a summarization algorithm. This algorithm efficiently extracts important information and generates a summary.
[0354] Step 4:
[0355] The server automatically inputs the generated summary information into the questionnaire according to a standard format. Here, the information is accurately mapped to a pre-configured template.
[0356] Step 5:
[0357] Simultaneously, the device's emotion engine analyzes the conversation between the user and the person being cared for, recognizing their emotional state from their tone of voice and the flow of the conversation. Emotional data is accumulated in real time.
[0358] Step 6:
[0359] The server verifies the submitted questionnaires to check for any omissions or inconsistencies. During this process, it utilizes data from the emotion engine, focusing particularly on areas where emotional responses were strong.
[0360] Step 7:
[0361] If any omissions or inconsistencies are detected, the server will consider the emotional context and generate appropriate additional questions. The generated questions will be notified to the user via the terminal.
[0362] Step 8:
[0363] The user follows the instructions on the device, asks additional questions to the person being cared for, and records their responses. The newly obtained audio data is processed according to the aforementioned flow and reflected in the questionnaire.
[0364] Step 9:
[0365] Finally, the server verifies the content and securely stores the complete questionnaire in cloud storage or a database. This improves the accuracy of the survey records.
[0366] (Example 2)
[0367] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0368] The aim is to streamline the process of home visit surveys, from acquiring audio data to inputting information and analyzing emotional states, thereby ensuring accuracy and smooth progress of the surveys. Furthermore, the goal is to provide a system that reduces the psychological burden on care recipients and surveyors during the survey, enabling the acquisition of highly reliable data.
[0369] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0370] In this invention, the server includes an acquisition means for acquiring audio during a visit, a conversion means for converting audio into text information, a summarization means for extracting and summarizing important information, an emotion analysis means for evaluating emotional states through audio data analysis, and a generation means for verifying input data and generating additional inquiries. This enables efficient and accurate data acquisition during on-site surveys, as well as flexible survey progress that takes emotional states into consideration.
[0371] "Acquisition method" refers to a function for collecting audio data during visits while suppressing environmental noise.
[0372] The "conversion means" is a function that converts acquired speech into text information, and it uses a highly accurate language model.
[0373] A "summarization tool" is a function that extracts important information from text and converts it into a summarized format.
[0374] "Input method" refers to a function that automatically inputs summarized information into a standardized information format.
[0375] "Emotion analysis means" refers to a function that evaluates the emotional state of the user and the person they are talking to based on voice data.
[0376] A "detection mechanism" is a function that verifies input data, checks for omissions or inconsistencies, and adjusts the information.
[0377] A "generation mechanism" is a function that generates additional inquiries based on detected problems and missing information, and presents them to the interlocutor.
[0378] To implement this invention, a system is constructed that streamlines the on-site survey process and improves its accuracy by using the following hardware and software.
[0379] First, the device is equipped with an audio acquisition module and uses noise cancellation technology to acquire clear audio data. High-performance audio capture devices and noise reduction software are used for audio acquisition. For example, a microphone built into a smart device carried by the researcher and a signal processing algorithm that suppresses ambient noise could be considered.
[0380] Next, the device converts the audio data into text information using a conversion method. A cloud-based speech recognition API is used for speech recognition, and a multilingual language model handles dialects and intonation differences. For example, the Google Cloud Speech-to-Text API is used to convert audio data into text.
[0381] The server analyzes the received text information using a summarization mechanism, extracting and summarizing important information. Using natural language processing techniques, a specific algorithm automatically summarizes the key points of the information. For example, an algorithm that analyzes word frequency and relationships is used for information extraction.
[0382] On the other hand, the emotion analysis system analyzes voice data and evaluates the emotions of the user and the person they are speaking with. An emotion recognition algorithm is implemented to analyze the emotional state in real time based on the tone and content of the voice. This allows for the evaluation of changes in the emotions of the person being surveyed during the investigation.
[0383] The server uses detection methods to verify the input data and compare it to a template to check for any omissions or inconsistencies. By also considering the results of sentiment analysis, it is possible to adjust the questions if the user is experiencing stress.
[0384] Finally, the server generates additional queries to complement the detected problems using a generation mechanism. Using a generation AI model, it creates new questions based on appropriate prompts and presents them to the user via the terminal. An example of a prompt might be, "In a home visit, please describe how to analyze elderly people's feelings about food and identify their specific preferences."
[0385] This system improves the accuracy of on-site surveys, reduces the psychological burden on both the user and the interviewee, and enables smooth and reliable surveys.
[0386] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0387] Step 1:
[0388] The device activates its voice acquisition module at the start of the on-site survey and records voice data in real time while suppressing ambient noise using noise cancellation technology. The input is the conversation with the survey subject, and the output is a clear audio file. Specifically, the device's recording function is activated, and noise filtering is performed in the background.
[0389] Step 2:
[0390] The device converts recorded audio into text information using a conversion mechanism. The input is an audio file, and the output is text. Using a speech recognition API, a language model analyzes the audio waveform to convert it into text. Specifically, the audio is sent to a cloud service, and each audio segment is mapped to its corresponding text.
[0391] Step 3:
[0392] The server receives the converted text information, extracts important information using a summarization tool, and performs a summary. The input is text, and the output is summarized information. Natural language processing technology is used to analyze keywords and themes from the text and perform a summary. Specifically, a text analysis algorithm scans the content and selects sentences of high importance.
[0393] Step 4:
[0394] The server automatically inputs summarized information into a standardized information format. The input is summarized information, and the output is a standardized survey record. A database management system is used, and the summarization results are stored in the database according to the fields. Specifically, the information is automatically entered according to a predetermined format and registered as a survey record.
[0395] Step 5:
[0396] The emotion analysis tool analyzes the audio data and evaluates the emotional state of the user and the person they are speaking with. The input is the audio data, and the output is a score indicating the emotional state. The emotion identification algorithm analyzes the tone and pitch of the voice to identify changes in emotion. Specifically, the audio feature extraction tool scans the audio and assigns emotion labels.
[0397] Step 6:
[0398] The server uses detection mechanisms to check for missing or inconsistent data in the input and adjusts the data to take sentiment into account. The input consists of standardized survey records and sentiment scores, and the output is the corrected survey records. If inconsistencies are detected, the system re-evaluates the data and generates adjustment proposals. Specifically, the algorithm identifies the locations of missing data and applies corrective processing as needed.
[0399] Step 7:
[0400] The server uses a generation mechanism to generate additional inquiries regarding detected problems and notifies the user via the terminal. The input is inconsistent data and sentiment information, and the output is the additional inquiries. A generation AI model creates prompts and generates the necessary additional questions. For example, if there is an ambiguous answer, a question such as "What specific dishes do you like?" is automatically generated and suggested to the user via the terminal.
[0401] (Application Example 2)
[0402] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0403] In many physical stores, communication between customers and service staff significantly impacts the quality of service. However, it is difficult for service staff to accurately grasp the emotions and needs of all customers in real time, leading to inconsistencies in service quality. Furthermore, there are insufficient resources available to improve customer satisfaction by appropriately analyzing and addressing customer emotions and needs.
[0404] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0405] In this invention, the server includes a collection means for acquiring audio during a visit, a conversion means for converting the collected audio into text information, and an emotion analysis means for analyzing the audio and evaluating the emotional state. This makes it possible to improve the quality of communication between customers and service staff in physical stores, understand the customer's emotional state, and improve the response.
[0406] "Collection means" refers to a mechanism for acquiring voice data during visits.
[0407] "Conversion means" refers to technology for converting acquired audio data into text information.
[0408] A "summarization method" is a technique that extracts and summarizes important information from converted textual information.
[0409] "Input method" refers to technology that automatically inputs summarized information into a standard format.
[0410] A "detection means" is a mechanism for detecting omissions or inconsistencies in the input data.
[0411] "Generation means" refers to techniques for generating additional questions based on detected problems.
[0412] "Emotional analysis methods" refer to technologies that analyze voice to evaluate emotional states.
[0413] "Suggestion methods" refer to methods for improving customer service based on information obtained through emotion analysis.
[0414] The embodiment for carrying out the invention is configured as follows as a method for realizing a system to efficiently improve customer service in physical stores. This system has the function of collecting voice in real time, converting it into text information, and summarizing it. In addition, it analyzes the emotional state of customers from the voice data and uses that information to make suggestions for improving customer service.
[0415] The server uses speech recognition software to collect conversations between customers and service staff during visits. This collection method utilizes a recording device employing noise-canceling technology. Next, the collected audio data is converted into text information using natural language processing. At this stage, dialects and intonation are taken into consideration to generate accurate text data.
[0416] The converted textual information is summarized by a summarization algorithm, which extracts the most important information. This summarized information is then automatically entered according to a predetermined format. After input, the server uses detection means to check for any missing or inconsistent data, and generates additional questions as needed using generation means.
[0417] Furthermore, the server utilizes emotion analysis techniques to analyze text derived from the audio data and assess the customer's emotional state. This analysis employs a generative AI model that captures emotional changes based on voice tone and context.
[0418] For example, if a customer expresses positive feelings towards a product, the server can use this information to make further suggestions to the customer service representative, offering more helpful service. An example of a prompt would be, "Please provide sentiment analysis and summary of this conversation: 'I liked this product, but I have one more question.'"
[0419] This will improve the quality of customer service at physical stores and increase customer satisfaction.
[0420] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0421] Step 1:
[0422] The terminal activates the voice acquisition module and collects conversations between customers and service staff in real time during visits. Using noise cancellation technology, it acquires clear voice data while minimizing ambient noise. An audio signal is provided as input, and noise-reduced voice data is obtained as output.
[0423] Step 2:
[0424] The server processes the collected audio data using speech recognition software and converts it into text information. Speech recognition technology is used to convert audio information into text data. At this stage, the input is denoised audio data, and the output is recognized text data.
[0425] Step 3:
[0426] The server processes the converted character information using a summarization algorithm, extracting important information and creating a summary. This process uses natural language processing techniques to extract key points and shorten the text. The input is recognized character data, and the output is summarized text information.
[0427] Step 4:
[0428] The server automatically inputs summarized information into a pre-configured format. This format has a fixed data structure, and the data is automatically entered according to it. The input is summarized text information, and the output is a pre-filled datasheet according to the format.
[0429] Step 5:
[0430] The server analyzes the input data using detection methods to check for any missing or inconsistent data. The purpose of this process is to verify data integrity and identify any problems. The input is formatted data, and the output is the result of data verification.
[0431] Step 6:
[0432] Based on the problems detected by the server, the generation mechanism generates additional questions and notifies the user via the terminal. Here, questions are dynamically created if there are problems. The input is the verification result, and the output is a set of follow-up questions.
[0433] Step 7:
[0434] The server uses emotion analysis tools to evaluate the emotional state of text obtained from speech. A generative AI model is used to analyze voice tone and context. The input is textual information, and the output is emotional state data.
[0435] Step 8:
[0436] Based on the results of sentiment analysis, the server activates a suggestion mechanism to propose service improvements to the user. The results are returned as suggestions for the service to the user. The input is sentiment state data, and the output is improvement suggestions.
[0437] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0438] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0439] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0440] [Third Embodiment]
[0441] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0442] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0443] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0444] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0445] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0446] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0447] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0448] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0449] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0450] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0451] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0452] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0453] This invention relates to the construction of a system that enables the acquisition of audio data during on-site surveys and its automatic input into questionnaires. The core of the system consists of acquisition means, conversion means, summarization means, input means, detection means, and generation means. Embodiments of each means are described below.
[0454] First, once the home visit begins, the device starts acquiring audio data. At this time, the device is equipped with a high-sensitivity microphone and noise cancellation function to eliminate external noise while clearly recording the conversation between the care recipient and the staff member.
[0455] Next, the recorded audio data is converted into text in real time by a conversion mechanism within the device. The device uses an advanced speech recognition algorithm that can handle differences in dialect and intonation.
[0456] Once the text data is generated, the server activates a summarization mechanism. This mechanism extracts important information from the vast amount of text, selecting and shortening the necessary data. This creates the summarized data ready for input into the questionnaire.
[0457] The server then automatically inputs the summarized information into each field of the questionnaire. The input method uses data formatted in a standard format, arranging the information according to a template.
[0458] Furthermore, to ensure that all entered information is accurate, the server uses detection mechanisms to detect any missing or inconsistent information. This detection mechanism is achieved by utilizing AI models and comparing the data with existing databases.
[0459] Finally, if the detection means finds a problem, the server uses the generation means to form additional questions. These questions are sent to the terminal in real time to prompt the user for further confirmation.
[0460] For example, suppose during a home visit, a care recipient is asked, "What kind of exercise do you do every day?" and the answer is, "I take walks three times a week." The server summarizes this information and enters "Exercise: Walks three times a week" into the questionnaire. However, if there are any missing details regarding specific dates or other activities related to the "every day" question, the server generates additional questions to resolve the inconsistencies.
[0461] This system makes on-site surveys more efficient, allowing staff to improve the accuracy and reliability of survey results.
[0462] The following describes the processing flow.
[0463] Step 1:
[0464] The terminal activates its voice acquisition module at the start of the on-site survey and records the conversation in real time. The acquired voice data is temporarily stored on the local device.
[0465] Step 2:
[0466] After accumulating a certain amount of continuously acquired audio data in a buffer, the terminal activates speech recognition software to convert the audio into text. Noise reduction processing is applied during this conversion process to minimize misrecognition.
[0467] Step 3:
[0468] Once speech recognition is complete, the converted text data is sent to the server. The server inputs the received data into a summarization algorithm. Here, important information is extracted and a summarized text is generated.
[0469] Step 4:
[0470] Upon receiving the summarized text information, the server begins the automatic input process into the questionnaire template. During this process, the questionnaire is constructed sequentially, mapping each summarized piece of information to its corresponding survey item.
[0471] Step 5:
[0472] To verify the entered survey data, the server uses detection methods to compare it with existing data and templates, identifying any missing or inconsistent data. Anomalies are detected using AI models during this process.
[0473] Step 6:
[0474] Based on the identified omissions and inconsistencies, the server generates additional questions using a generation mechanism. The generated questions are sent to the terminal in a timely manner and provided as real-time feedback to the user.
[0475] Step 7:
[0476] Based on additional questions presented on the device, the user reconfirms information with the care recipient and records their responses. The newly acquired audio data is then recognized and processed again and reflected in the questionnaire.
[0477] Step 8:
[0478] Finally, the server fully verifies the contents of the questionnaire, and if there are no problems, it saves the completed questionnaire to secure storage. As a result, records of on-site surveys can be managed efficiently and accurately.
[0479] (Example 1)
[0480] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0481] In field surveys, it is necessary to improve the accuracy and efficiency of acquiring audio data accurately and efficiently, and then automatically inputting that data into questionnaires. Furthermore, a challenge is to enhance the reliability of survey results by reliably detecting missing or inconsistent information and quickly acquiring additional information.
[0482] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0483] In this invention, the server includes means for acquiring audio data during a visit, means for converting the acquired audio data into text information, and means for summarizing the converted text information and extracting important information. This enables efficient acquisition of audio data during on-site surveys and automatic input of accurate information based on that data.
[0484] A "device for acquiring audio data during a visit" refers to a terminal equipped with a high-sensitivity microphone and noise cancellation function to effectively eliminate external noise and record conversations clearly.
[0485] A "device that converts acquired audio data into text information" refers to a terminal equipped with a speech recognition algorithm that utilizes a generative AI model to convert audio data into accurate text data in real time.
[0486] A "device for summarizing converted text information and extracting important information" refers to a server that automatically extracts important information from generated text data and shortens it to suit the purpose of the investigation.
[0487] A "device that automatically inputs data into a standardized format" refers to a server equipped with data entry capabilities to neatly arrange summarized information onto a questionnaire based on a template.
[0488] A "device for detecting omissions and inconsistencies" refers to a server equipped with an AI model that identifies problems by comparing automatically entered data with existing databases to verify its integrity.
[0489] A "device for generating questions to obtain additional necessary information" refers to a server that, based on detected omissions and inconsistencies, generates new questions in real time as needed and sends them to the terminal.
[0490] This invention provides a system for streamlining the automatic acquisition of voice data and automatic input into questionnaires during on-site surveys. A specific embodiment of this system is described below.
[0491] Once the home visit begins, the device uses a high-sensitivity microphone and noise cancellation function to clearly record conversations between the care recipient and staff. This makes it possible to obtain accurate audio data while eliminating external noise.
[0492] Next, to convert the acquired audio data into text data in real time, the device uses an advanced speech recognition algorithm that applies a generative AI model. This algorithm takes into account dialects and the speaker's intonation, enabling it to generate accurate text information.
[0493] The generated text data is sent to a server where it is summarized. The server uses natural language processing techniques to extract important information from the large amount of text, narrowing down the necessary data. This removes redundant information and ensures that the summarized data necessary for the investigation is secured.
[0494] Next, the server automatically enters the summarized information into each item of the questionnaire based on a standardized format. This input process is designed to ensure that the data is precisely placed in the correct location, increasing the efficiency of automated entry.
[0495] To detect missing or inconsistent data, the server utilizes detection mechanisms. These mechanisms compare the data with existing databases and use AI models to verify its accuracy. This improves the reliability of the investigation results.
[0496] Finally, if any omissions or inconsistencies are found in the information, the server uses a generation mechanism to formulate additional questions and send them to the terminal. This allows the user to immediately verify and provide additional information.
[0497] For example, suppose during a home visit, a care recipient is asked, "What kind of exercise do you do every day?" and the answer is, "I take walks three times a week." The server summarizes this information as "Exercise: Walks three times a week" and enters it into the questionnaire. As a result, the efficiency and accuracy of the survey are improved.
[0498] An example of a prompt to input into a generative AI model is, "Please tell me how to convert audio data obtained from a field survey into text, summarize the important information, and input it into a survey form."
[0499] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0500] Step 1:
[0501] The terminal begins acquiring audio data simultaneously with the start of the home visit. A high-sensitivity microphone captures external conversational audio as input. By utilizing noise cancellation technology to remove unwanted background noise, clear conversation data between the care recipient and staff is output. This audio data is a crucial source of information that forms the basis of processing.
[0502] Step 2:
[0503] The device activates its built-in speech recognition software to convert the acquired audio data into text. Audio data is used as input, and advanced algorithms from a generative AI model are applied to handle speech characteristics such as dialects and intonation. The output of this process is accurate, structured text data.
[0504] Step 3:
[0505] The server receives text data and uses summarization techniques. Upon receiving text data, it utilizes natural language processing techniques to extract important information from the vast amount of data. At this stage, it removes extraneous data and outputs a shortened summary necessary for the investigation.
[0506] Step 4:
[0507] The server automatically inputs the summarized information according to the standard questionnaire format. In this process, the questionnaire template is used as input, and the summarized information is organized so that it corresponds to each item. The output is properly completed questionnaire data.
[0508] Step 5:
[0509] The server activates detection mechanisms to verify the integrity of the input data. The input consists of completed questionnaires, and by combining an AI model and a database, it identifies omissions and inconsistencies. If problems are detected during this process, the findings are provided as output.
[0510] Step 6:
[0511] When a problem is detected, the server uses a generation mechanism to form any necessary additional questions. The output from the detection mechanism becomes the input, and the generated questions are output. These questions are sent to the terminal in real time, requesting additional information from the user and ensuring data integrity with the newly obtained information.
[0512] (Application Example 1)
[0513] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0514] Traditional in-person surveys faced challenges in acquiring voice data and efficiently inputting that data, as well as difficulty in understanding customer needs in real time and effectively recommending products. This hindered improvements in survey accuracy and customer experience.
[0515] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0516] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting audio into text information, and a recommendation means for acquiring conversations with customers and making product recommendations. This enables improved accuracy of on-site surveys and an enhanced customer experience.
[0517] "Acquisition means" refers to a device or system for accurately acquiring audio data during a visit, while taking ambient noise into consideration.
[0518] "Conversion means" refers to the process or device for converting acquired audio data into a different format, specifically textual information.
[0519] A "summarization tool" is a process or system for extracting necessary information from converted textual information and displaying it in a shortened form.
[0520] An "input method" is a system for automatically inputting summarized information according to a standardized format.
[0521] A "detection means" is a system that has the function of comparing input data with a template to discover missing or inconsistent information.
[0522] A "recommendation method" is a system that selects and presents the most suitable products based on information obtained from conversations with customers.
[0523] A "generation means" is a process or system that has the function of generating additional questions based on the detected problem.
[0524] This invention relates to the implementation of a system for inputting information from voice data obtained through on-site surveys and in physical stores, as well as a product recommendation system. Efficient and highly accurate information processing is achieved through the cooperation of a server and terminals.
[0525] The device records conversations with customers during visits or within stores using a high-sensitivity microphone and obtains clear audio data by employing technology to remove background noise. The recorded audio is converted into text in real time using speech recognition software (e.g., Google Speech-to-Text API) on the device. This process also supports diverse speech patterns, including intonation and dialects.
[0526] Subsequently, the server receives the converted text information and summarizes it using a natural language processing engine (e.g., spaCy). It extracts key information and condenses it into the information that forms the basis for each item on the questionnaire and for product recommendations. The summarized information is then formatted into a standard format and automatically entered into the questionnaire and product database.
[0527] After input, the server activates a detection mechanism and uses an AI model (e.g., TensorFlow) to detect missing or inconsistent data. During this process, it compares the data with an existing database, generates additional questions as needed, and sends them to the terminal. The user can then use the terminal to answer these additional questions for further verification.
[0528] For example, in a physical store, if a customer asks, "I'm having a barbecue with friends next Sunday, what do I need?", the server analyzes the conversation and recommends products such as "barbecue meat" and "seasoning sets," and also provides relevant discount information.
[0529] Examples of prompts for a generative AI model:
[0530] "The customer is looking for barbecue-related products. Please generate a list of required items and a list of recommended products."
[0531] In this way, it is possible to improve the efficiency and accuracy of on-site surveys, as well as enhance the customer experience at physical stores.
[0532] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0533] Step 1:
[0534] The device uses a high-sensitivity microphone to capture ambient sound and noise cancellation to obtain clear audio data. It receives conversations between customers and staff as input and outputs a clear audio file, which is then used for subsequent processing.
[0535] Step 2:
[0536] The device converts acquired audio data into text in real time using speech recognition software (e.g., Google Speech-to-Text API). It analyzes the input audio data and outputs the corresponding text data. This process also takes into account differences in intonation and dialect.
[0537] Step 3:
[0538] The server receives text data sent from the terminal and uses a natural language processing engine (e.g., spaCy) to perform summarization. It extracts important information and removes unnecessary parts to output summarized text. This summarized data is then used in the next step.
[0539] Step 4:
[0540] The server automatically formats the summarized text information into a standard format and inputs it into the questionnaire or product database. The input summarized information is then formatted, and formatted data is obtained as output. This data is then adapted to a known template.
[0541] Step 5:
[0542] The server uses an AI model (e.g., TensorFlow) to detect missing or inconsistent data in the input. This AI model compares the input data against an existing database to verify its integrity. If problems are found, the detection results are provided as output.
[0543] Step 6:
[0544] The server generates additional questions based on the detected issues and sends them to the terminal. This generation AI model is used to create specific question prompts. The additional questions are then provided to the user as output.
[0545] Step 7:
[0546] The user follows the additional questions displayed on the terminal and enters their responses. The terminal receives the entered responses and sends the data to the server. This prepares the server to supply highly accurate data as output for the next processing step.
[0547] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0548] This invention relates to a system that not only acquires voice data during on-site surveys and automatically inputs it into questionnaires, but also incorporates an emotion engine that recognizes the user's emotional state. This system includes acquisition means, conversion means, summarization means, input means, detection means, generation means, and an emotion engine. Detailed embodiments of each means and the emotion engine are described below.
[0549] To begin, the device activates its voice acquisition module at the start of the visit and records the conversation in real time. At this stage, noise cancellation technology is used to minimize ambient noise while acquiring clear voice data.
[0550] The recorded audio is immediately converted into text information by the device's built-in conversion mechanism. This conversion process utilizes diverse language models to accommodate dialects and intonation differences, achieving highly accurate text conversion.
[0551] Next, the server receives the converted text and uses a summarization tool to extract and summarize the important information. The generated summary information is automatically entered into the questionnaire based on a standardized format.
[0552] A key feature of this system is its emotion engine, which simultaneously analyzes voice data and evaluates the emotional state of both the user and the person being cared for. The emotion engine analyzes voice tone, speaking style, and context to grasp changes in emotional state in real time.
[0553] The server uses detection methods to check for any missing or inconsistent data in the entered questionnaire. During this process, it also utilizes emotional information obtained by the emotion engine to adjust the questions based on whether the user is experiencing stress or based on their responses.
[0554] Furthermore, if the detection means identifies any omissions or inconsistencies in the input data, the server uses the generation means to generate additional questions and notifies the user via the terminal.
[0555] For example, if a person receiving care gives a vague answer to the question, "Do you enjoy the food?", the server will consider their emotional state and suggest follow-up questions such as, "What kind of food do you particularly like?"
[0556] This system not only improves the accuracy of surveys but also reduces the psychological burden on users, supporting the smooth and reliable conduct of surveys.
[0557] The following describes the processing flow.
[0558] Step 1:
[0559] The device activates its voice acquisition function as soon as the on-site survey begins, recording the conversation clearly. Recording takes place in the background, and ambient noise is removed using the noise cancellation function.
[0560] Step 2:
[0561] The device transfers the acquired audio data to the speech recognition system in real time, where it is converted into text. The converted text is then saved in a format that can be processed immediately.
[0562] Step 3:
[0563] The server receives text data sent from the terminal and applies a summarization algorithm. This algorithm efficiently extracts important information and generates a summary.
[0564] Step 4:
[0565] The server automatically inputs the generated summary information into the questionnaire according to a standard format. Here, the information is accurately mapped to a pre-configured template.
[0566] Step 5:
[0567] Simultaneously, the device's emotion engine analyzes the conversation between the user and the person being cared for, recognizing their emotional state from their tone of voice and the flow of the conversation. Emotional data is accumulated in real time.
[0568] Step 6:
[0569] The server verifies the submitted questionnaires to check for any omissions or inconsistencies. During this process, it utilizes data from the emotion engine, focusing particularly on areas where emotional responses were strong.
[0570] Step 7:
[0571] If any omissions or inconsistencies are detected, the server will consider the emotional context and generate appropriate additional questions. The generated questions will be notified to the user via the terminal.
[0572] Step 8:
[0573] The user follows the instructions on the device, asks additional questions to the person being cared for, and records their responses. The newly obtained audio data is processed according to the aforementioned flow and reflected in the questionnaire.
[0574] Step 9:
[0575] Finally, the server verifies the content and securely stores the complete questionnaire in cloud storage or a database. This improves the accuracy of the survey records.
[0576] (Example 2)
[0577] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0578] The aim is to streamline the process of home visit surveys, from acquiring audio data to inputting information and analyzing emotional states, thereby ensuring accuracy and smooth progress of the surveys. Furthermore, the goal is to provide a system that reduces the psychological burden on care recipients and surveyors during the survey, enabling the acquisition of highly reliable data.
[0579] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0580] In this invention, the server includes an acquisition means for acquiring audio during a visit, a conversion means for converting audio into text information, a summarization means for extracting and summarizing important information, an emotion analysis means for evaluating emotional states through audio data analysis, and a generation means for verifying input data and generating additional inquiries. This enables efficient and accurate data acquisition during on-site surveys, as well as flexible survey progress that takes emotional states into consideration.
[0581] "Acquisition method" refers to a function for collecting audio data during visits while suppressing environmental noise.
[0582] The "conversion means" is a function that converts acquired speech into text information, and it uses a highly accurate language model.
[0583] A "summarization tool" is a function that extracts important information from text and converts it into a summarized format.
[0584] "Input method" refers to a function that automatically inputs summarized information into a standardized information format.
[0585] "Emotion analysis means" refers to a function that evaluates the emotional state of the user and the person they are talking to based on voice data.
[0586] A "detection mechanism" is a function that verifies input data, checks for omissions or inconsistencies, and adjusts the information.
[0587] A "generation mechanism" is a function that generates additional inquiries based on detected problems and missing information, and presents them to the interlocutor.
[0588] To implement this invention, a system is constructed that streamlines the on-site survey process and improves its accuracy by using the following hardware and software.
[0589] First, the device is equipped with an audio acquisition module and uses noise cancellation technology to acquire clear audio data. High-performance audio capture devices and noise reduction software are used for audio acquisition. For example, a microphone built into a smart device carried by the researcher and a signal processing algorithm that suppresses ambient noise could be considered.
[0590] Next, the device converts the audio data into text information using a conversion method. A cloud-based speech recognition API is used for speech recognition, and a multilingual language model handles dialects and intonation differences. For example, the Google Cloud Speech-to-Text API is used to convert audio data into text.
[0591] The server analyzes the received text information using a summarization mechanism, extracting and summarizing important information. Using natural language processing techniques, a specific algorithm automatically summarizes the key points of the information. For example, an algorithm that analyzes word frequency and relationships is used for information extraction.
[0592] On the other hand, the emotion analysis system analyzes voice data and evaluates the emotions of the user and the person they are speaking with. An emotion recognition algorithm is implemented to analyze the emotional state in real time based on the tone and content of the voice. This allows for the evaluation of changes in the emotions of the person being surveyed during the investigation.
[0593] The server uses detection methods to verify the input data and compare it to a template to check for any omissions or inconsistencies. By also considering the results of sentiment analysis, it is possible to adjust the questions if the user is experiencing stress.
[0594] Finally, the server generates additional queries to complement the detected problems using a generation mechanism. Using a generation AI model, it creates new questions based on appropriate prompts and presents them to the user via the terminal. An example of a prompt might be, "In a home visit, please describe how to analyze elderly people's feelings about food and identify their specific preferences."
[0595] This system improves the accuracy of on-site surveys, reduces the psychological burden on both the user and the interviewee, and enables smooth and reliable surveys.
[0596] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0597] Step 1:
[0598] The device activates its voice acquisition module at the start of the on-site survey and records voice data in real time while suppressing ambient noise using noise cancellation technology. The input is the conversation with the survey subject, and the output is a clear audio file. Specifically, the device's recording function is activated, and noise filtering is performed in the background.
[0599] Step 2:
[0600] The device converts recorded audio into text information using a conversion mechanism. The input is an audio file, and the output is text. Using a speech recognition API, a language model analyzes the audio waveform to convert it into text. Specifically, the audio is sent to a cloud service, and each audio segment is mapped to its corresponding text.
[0601] Step 3:
[0602] The server receives the converted text information, extracts important information using a summarization tool, and performs a summary. The input is text, and the output is summarized information. Natural language processing technology is used to analyze keywords and themes from the text and perform a summary. Specifically, a text analysis algorithm scans the content and selects sentences of high importance.
[0603] Step 4:
[0604] The server automatically inputs summarized information into a standardized information format. The input is summarized information, and the output is a standardized survey record. A database management system is used, and the summarization results are stored in the database according to the fields. Specifically, the information is automatically entered according to a predetermined format and registered as a survey record.
[0605] Step 5:
[0606] The emotion analysis tool analyzes the audio data and evaluates the emotional state of the user and the person they are speaking with. The input is the audio data, and the output is a score indicating the emotional state. The emotion identification algorithm analyzes the tone and pitch of the voice to identify changes in emotion. Specifically, the audio feature extraction tool scans the audio and assigns emotion labels.
[0607] Step 6:
[0608] The server uses detection mechanisms to check for missing or inconsistent data in the input and adjusts the data to take sentiment into account. The input consists of standardized survey records and sentiment scores, and the output is the corrected survey records. If inconsistencies are detected, the system re-evaluates the data and generates adjustment proposals. Specifically, the algorithm identifies the locations of missing data and applies corrective processing as needed.
[0609] Step 7:
[0610] The server uses a generation mechanism to generate additional inquiries regarding detected problems and notifies the user via the terminal. The input is inconsistent data and sentiment information, and the output is the additional inquiries. A generation AI model creates prompts and generates the necessary additional questions. For example, if there is an ambiguous answer, a question such as "What specific dishes do you like?" is automatically generated and suggested to the user via the terminal.
[0611] (Application Example 2)
[0612] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0613] In many physical stores, communication between customers and service staff significantly impacts the quality of service. However, it is difficult for service staff to accurately grasp the emotions and needs of all customers in real time, leading to inconsistencies in service quality. Furthermore, there are insufficient resources available to improve customer satisfaction by appropriately analyzing and addressing customer emotions and needs.
[0614] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0615] In this invention, the server includes a collection means for acquiring audio during a visit, a conversion means for converting the collected audio into text information, and an emotion analysis means for analyzing the audio and evaluating the emotional state. This makes it possible to improve the quality of communication between customers and service staff in physical stores, understand the customer's emotional state, and improve the response.
[0616] "Collection means" refers to a mechanism for acquiring voice data during visits.
[0617] "Conversion means" refers to technology for converting acquired audio data into text information.
[0618] A "summarization method" is a technique that extracts and summarizes important information from converted textual information.
[0619] "Input method" refers to technology that automatically inputs summarized information into a standard format.
[0620] A "detection means" is a mechanism for detecting omissions or inconsistencies in the input data.
[0621] "Generation means" refers to techniques for generating additional questions based on detected problems.
[0622] "Emotional analysis methods" refer to technologies that analyze voice to evaluate emotional states.
[0623] "Suggestion methods" refer to methods for improving customer service based on information obtained through emotion analysis.
[0624] The embodiment for carrying out the invention is configured as follows as a method for realizing a system to efficiently improve customer service in physical stores. This system has the function of collecting voice in real time, converting it into text information, and summarizing it. In addition, it analyzes the emotional state of customers from the voice data and uses that information to make suggestions for improving customer service.
[0625] The server uses speech recognition software to collect conversations between customers and service staff during visits. This collection method utilizes a recording device employing noise-canceling technology. Next, the collected audio data is converted into text information using natural language processing. At this stage, dialects and intonation are taken into consideration to generate accurate text data.
[0626] The converted textual information is summarized by a summarization algorithm, which extracts the most important information. This summarized information is then automatically entered according to a predetermined format. After input, the server uses detection means to check for any missing or inconsistent data, and generates additional questions as needed using generation means.
[0627] Furthermore, the server utilizes emotion analysis techniques to analyze text derived from the audio data and assess the customer's emotional state. This analysis employs a generative AI model that captures emotional changes based on voice tone and context.
[0628] For example, if a customer expresses positive feelings towards a product, the server can use this information to make further suggestions to the customer service representative, offering more helpful service. An example of a prompt would be, "Please provide sentiment analysis and summary of this conversation: 'I liked this product, but I have one more question.'"
[0629] This will improve the quality of customer service at physical stores and increase customer satisfaction.
[0630] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0631] Step 1:
[0632] The terminal activates the voice acquisition module and collects conversations between customers and service staff in real time during visits. Using noise cancellation technology, it acquires clear voice data while minimizing ambient noise. An audio signal is provided as input, and noise-reduced voice data is obtained as output.
[0633] Step 2:
[0634] The server processes the collected audio data using speech recognition software and converts it into text information. Speech recognition technology is used to convert audio information into text data. At this stage, the input is denoised audio data, and the output is recognized text data.
[0635] Step 3:
[0636] The server processes the converted character information using a summarization algorithm, extracting important information and creating a summary. This process uses natural language processing techniques to extract key points and shorten the text. The input is recognized character data, and the output is summarized text information.
[0637] Step 4:
[0638] The server automatically inputs summarized information into a pre-configured format. This format has a fixed data structure, and the data is automatically entered according to it. The input is summarized text information, and the output is a pre-filled datasheet according to the format.
[0639] Step 5:
[0640] The server analyzes the input data using detection methods to check for any missing or inconsistent data. The purpose of this process is to verify data integrity and identify any problems. The input is formatted data, and the output is the result of data verification.
[0641] Step 6:
[0642] Based on the problems detected by the server, the generation mechanism generates additional questions and notifies the user via the terminal. Here, questions are dynamically created if there are problems. The input is the verification result, and the output is a set of follow-up questions.
[0643] Step 7:
[0644] The server uses emotion analysis tools to evaluate the emotional state of text obtained from speech. A generative AI model is used to analyze voice tone and context. The input is textual information, and the output is emotional state data.
[0645] Step 8:
[0646] Based on the results of sentiment analysis, the server activates a suggestion mechanism to propose service improvements to the user. The results are returned as suggestions for the service to the user. The input is sentiment state data, and the output is improvement suggestions.
[0647] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0648] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0649] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0650] [Fourth Embodiment]
[0651] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0652] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0653] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0654] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0655] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0656] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0657] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0658] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0659] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0660] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0661] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0662] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0663] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0664] This invention relates to the construction of a system that enables the acquisition of audio data during on-site surveys and its automatic input into questionnaires. The core of the system consists of acquisition means, conversion means, summarization means, input means, detection means, and generation means. Embodiments of each means are described below.
[0665] First, once the home visit begins, the device starts acquiring audio data. At this time, the device is equipped with a high-sensitivity microphone and noise cancellation function to eliminate external noise while clearly recording the conversation between the care recipient and the staff member.
[0666] Next, the recorded audio data is converted into text in real time by a conversion mechanism within the device. The device uses an advanced speech recognition algorithm that can handle differences in dialect and intonation.
[0667] Once the text data is generated, the server activates a summarization mechanism. This mechanism extracts important information from the vast amount of text, selecting and shortening the necessary data. This creates the summarized data ready for input into the questionnaire.
[0668] The server then automatically inputs the summarized information into each field of the questionnaire. The input method uses data formatted in a standard format, arranging the information according to a template.
[0669] Furthermore, to ensure that all entered information is accurate, the server uses detection mechanisms to detect any missing or inconsistent information. This detection mechanism is achieved by utilizing AI models and comparing the data with existing databases.
[0670] Finally, if the detection means finds a problem, the server uses the generation means to form additional questions. These questions are sent to the terminal in real time to prompt the user for further confirmation.
[0671] For example, suppose during a home visit, a care recipient is asked, "What kind of exercise do you do every day?" and the answer is, "I take walks three times a week." The server summarizes this information and enters "Exercise: Walks three times a week" into the questionnaire. However, if there are any missing details regarding specific dates or other activities related to the "every day" question, the server generates additional questions to resolve the inconsistencies.
[0672] This system makes on-site surveys more efficient, allowing staff to improve the accuracy and reliability of survey results.
[0673] The following describes the processing flow.
[0674] Step 1:
[0675] The terminal activates its voice acquisition module at the start of the on-site survey and records the conversation in real time. The acquired voice data is temporarily stored on the local device.
[0676] Step 2:
[0677] After accumulating a certain amount of continuously acquired audio data in a buffer, the terminal activates speech recognition software to convert the audio into text. Noise reduction processing is applied during this conversion process to minimize misrecognition.
[0678] Step 3:
[0679] Once speech recognition is complete, the converted text data is sent to the server. The server inputs the received data into a summarization algorithm. Here, important information is extracted and a summarized text is generated.
[0680] Step 4:
[0681] Upon receiving the summarized text information, the server begins the automatic input process into the questionnaire template. During this process, the questionnaire is constructed sequentially, mapping each summarized piece of information to its corresponding survey item.
[0682] Step 5:
[0683] To verify the entered survey data, the server uses detection methods to compare it with existing data and templates, identifying any missing or inconsistent data. Anomalies are detected using AI models during this process.
[0684] Step 6:
[0685] Based on the identified omissions and inconsistencies, the server generates additional questions using a generation mechanism. The generated questions are sent to the terminal in a timely manner and provided as real-time feedback to the user.
[0686] Step 7:
[0687] Based on additional questions presented on the device, the user reconfirms information with the care recipient and records their responses. The newly acquired audio data is then recognized and processed again and reflected in the questionnaire.
[0688] Step 8:
[0689] Finally, the server fully verifies the contents of the questionnaire, and if there are no problems, it saves the completed questionnaire to secure storage. As a result, records of on-site surveys can be managed efficiently and accurately.
[0690] (Example 1)
[0691] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0692] In field surveys, it is necessary to improve the accuracy and efficiency of acquiring audio data accurately and efficiently, and then automatically inputting that data into questionnaires. Furthermore, a challenge is to enhance the reliability of survey results by reliably detecting missing or inconsistent information and quickly acquiring additional information.
[0693] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0694] In this invention, the server includes means for acquiring audio data during a visit, means for converting the acquired audio data into text information, and means for summarizing the converted text information and extracting important information. This enables efficient acquisition of audio data during on-site surveys and automatic input of accurate information based on that data.
[0695] A "device for acquiring audio data during a visit" refers to a terminal equipped with a high-sensitivity microphone and noise cancellation function to effectively eliminate external noise and record conversations clearly.
[0696] A "device that converts acquired audio data into text information" refers to a terminal equipped with a speech recognition algorithm that utilizes a generative AI model to convert audio data into accurate text data in real time.
[0697] A "device for summarizing converted text information and extracting important information" refers to a server that automatically extracts important information from generated text data and shortens it to suit the purpose of the investigation.
[0698] A "device that automatically inputs data into a standardized format" refers to a server equipped with data entry capabilities to neatly arrange summarized information onto a questionnaire based on a template.
[0699] A "device for detecting omissions and inconsistencies" refers to a server equipped with an AI model that identifies problems by comparing automatically entered data with existing databases to verify its integrity.
[0700] A "device for generating questions to obtain additional necessary information" refers to a server that, based on detected omissions and inconsistencies, generates new questions in real time as needed and sends them to the terminal.
[0701] This invention provides a system for streamlining the automatic acquisition of voice data and automatic input into questionnaires during on-site surveys. A specific embodiment of this system is described below.
[0702] Once the home visit begins, the device uses a high-sensitivity microphone and noise cancellation function to clearly record conversations between the care recipient and staff. This makes it possible to obtain accurate audio data while eliminating external noise.
[0703] Next, to convert the acquired audio data into text data in real time, the device uses an advanced speech recognition algorithm that applies a generative AI model. This algorithm takes into account dialects and the speaker's intonation, enabling it to generate accurate text information.
[0704] The generated text data is sent to a server where it is summarized. The server uses natural language processing techniques to extract important information from the large amount of text, narrowing down the necessary data. This removes redundant information and ensures that the summarized data necessary for the investigation is secured.
[0705] Next, the server automatically enters the summarized information into each item of the questionnaire based on a standardized format. This input process is designed to ensure that the data is precisely placed in the correct location, increasing the efficiency of automated entry.
[0706] To detect missing or inconsistent data, the server utilizes detection mechanisms. These mechanisms compare the data with existing databases and use AI models to verify its accuracy. This improves the reliability of the investigation results.
[0707] Finally, if any omissions or inconsistencies are found in the information, the server uses a generation mechanism to formulate additional questions and send them to the terminal. This allows the user to immediately verify and provide additional information.
[0708] For example, suppose during a home visit, a care recipient is asked, "What kind of exercise do you do every day?" and the answer is, "I take walks three times a week." The server summarizes this information as "Exercise: Walks three times a week" and enters it into the questionnaire. As a result, the efficiency and accuracy of the survey are improved.
[0709] An example of a prompt to input into a generative AI model is, "Please tell me how to convert audio data obtained from a field survey into text, summarize the important information, and input it into a survey form."
[0710] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0711] Step 1:
[0712] The terminal begins acquiring audio data simultaneously with the start of the home visit. A high-sensitivity microphone captures external conversational audio as input. By utilizing noise cancellation technology to remove unwanted background noise, clear conversation data between the care recipient and staff is output. This audio data is a crucial source of information that forms the basis of processing.
[0713] Step 2:
[0714] The device activates its built-in speech recognition software to convert the acquired audio data into text. Audio data is used as input, and advanced algorithms from a generative AI model are applied to handle speech characteristics such as dialects and intonation. The output of this process is accurate, structured text data.
[0715] Step 3:
[0716] The server receives text data and uses summarization techniques. Upon receiving text data, it utilizes natural language processing techniques to extract important information from the vast amount of data. At this stage, it removes extraneous data and outputs a shortened summary necessary for the investigation.
[0717] Step 4:
[0718] The server automatically inputs the summarized information according to the standard questionnaire format. In this process, the questionnaire template is used as input, and the summarized information is organized so that it corresponds to each item. The output is properly completed questionnaire data.
[0719] Step 5:
[0720] The server activates detection mechanisms to verify the integrity of the input data. The input consists of completed questionnaires, and by combining an AI model and a database, it identifies omissions and inconsistencies. If problems are detected during this process, the findings are provided as output.
[0721] Step 6:
[0722] When a problem is detected, the server uses a generation mechanism to form any necessary additional questions. The output from the detection mechanism becomes the input, and the generated questions are output. These questions are sent to the terminal in real time, requesting additional information from the user and ensuring data integrity with the newly obtained information.
[0723] (Application Example 1)
[0724] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0725] Traditional in-person surveys faced challenges in acquiring voice data and efficiently inputting that data, as well as difficulty in understanding customer needs in real time and effectively recommending products. This hindered improvements in survey accuracy and customer experience.
[0726] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0727] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting audio into text information, and a recommendation means for acquiring conversations with customers and making product recommendations. This enables improved accuracy of on-site surveys and an enhanced customer experience.
[0728] "Acquisition means" refers to a device or system for accurately acquiring audio data during a visit, while taking ambient noise into consideration.
[0729] "Conversion means" refers to the process or device for converting acquired audio data into a different format, specifically textual information.
[0730] A "summarization tool" is a process or system for extracting necessary information from converted textual information and displaying it in a shortened form.
[0731] An "input method" is a system for automatically inputting summarized information according to a standardized format.
[0732] A "detection means" is a system that has the function of comparing input data with a template to discover missing or inconsistent information.
[0733] A "recommendation method" is a system that selects and presents the most suitable products based on information obtained from conversations with customers.
[0734] A "generation means" is a process or system that has the function of generating additional questions based on the detected problem.
[0735] This invention relates to the implementation of a system for inputting information from voice data obtained through on-site surveys and in physical stores, as well as a product recommendation system. Efficient and highly accurate information processing is achieved through the cooperation of a server and terminals.
[0736] The device records conversations with customers during visits or within stores using a high-sensitivity microphone and obtains clear audio data by employing technology to remove background noise. The recorded audio is converted into text in real time using speech recognition software (e.g., Google Speech-to-Text API) on the device. This process also supports diverse speech patterns, including intonation and dialects.
[0737] Subsequently, the server receives the converted text information and summarizes it using a natural language processing engine (e.g., spaCy). It extracts key information and condenses it into the information that forms the basis for each item on the questionnaire and for product recommendations. The summarized information is then formatted into a standard format and automatically entered into the questionnaire and product database.
[0738] After input, the server activates a detection mechanism and uses an AI model (e.g., TensorFlow) to detect missing or inconsistent data. During this process, it compares the data with an existing database, generates additional questions as needed, and sends them to the terminal. The user can then use the terminal to answer these additional questions for further verification.
[0739] For example, in a physical store, if a customer asks, "I'm having a barbecue with friends next Sunday, what do I need?", the server analyzes the conversation and recommends products such as "barbecue meat" and "seasoning sets," and also provides relevant discount information.
[0740] Examples of prompts for a generative AI model:
[0741] "The customer is looking for barbecue-related products. Please generate a list of required items and a list of recommended products."
[0742] In this way, it is possible to improve the efficiency and accuracy of on-site surveys, as well as enhance the customer experience at physical stores.
[0743] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0744] Step 1:
[0745] The device uses a high-sensitivity microphone to capture ambient sounds and noise cancellation to obtain clear audio data. It receives conversations between customers and staff as input and outputs a clear audio file, which is then used for subsequent processing.
[0746] Step 2:
[0747] The device converts acquired audio data into text in real time using speech recognition software (e.g., Google Speech-to-Text API). It analyzes the input audio data and outputs the corresponding text data. This process also takes into account differences in intonation and dialect.
[0748] Step 3:
[0749] The server receives text data sent from the terminal and uses a natural language processing engine (e.g., spaCy) to perform summarization. It extracts important information and removes unnecessary parts to output summarized text. This summarized data is then used in the next step.
[0750] Step 4:
[0751] The server automatically formats the summarized text information into a standard format and inputs it into the questionnaire or product database. The input summarized information is then formatted, and formatted data is obtained as output. This data is then adapted to a known template.
[0752] Step 5:
[0753] The server uses an AI model (e.g., TensorFlow) to detect missing or inconsistent data in the input. This AI model compares the input data against an existing database to verify its integrity. If problems are found, the detection results are provided as output.
[0754] Step 6:
[0755] The server generates additional questions based on the detected issues and sends them to the terminal. This generation AI model is used to create specific question prompts. The additional questions are then provided to the user as output.
[0756] Step 7:
[0757] The user follows the additional questions displayed on the terminal and enters their responses. The terminal receives the entered responses and sends the data to the server. This prepares the server to supply highly accurate data as output for the next processing step.
[0758] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0759] This invention relates to a system that not only acquires voice data during on-site surveys and automatically inputs it into questionnaires, but also incorporates an emotion engine that recognizes the user's emotional state. This system includes acquisition means, conversion means, summarization means, input means, detection means, generation means, and an emotion engine. Detailed embodiments of each means and the emotion engine are described below.
[0760] To begin, the device activates its voice acquisition module at the start of the visit and records the conversation in real time. At this stage, noise cancellation technology is used to minimize ambient noise while acquiring clear voice data.
[0761] The recorded audio is immediately converted into text information by the device's built-in conversion mechanism. This conversion process utilizes diverse language models to accommodate dialects and intonation differences, achieving highly accurate text conversion.
[0762] Next, the server receives the converted text and uses a summarization tool to extract and summarize the important information. The generated summary information is automatically entered into the questionnaire based on a standardized format.
[0763] A key feature of this system is its emotion engine, which simultaneously analyzes voice data and evaluates the emotional state of both the user and the person being cared for. The emotion engine analyzes voice tone, speaking style, and context to grasp changes in emotional state in real time.
[0764] The server uses detection methods to check for any missing or inconsistent data in the entered questionnaire. During this process, it also utilizes emotional information obtained by the emotion engine to adjust the questions based on whether the user is experiencing stress or based on their responses.
[0765] Furthermore, if the detection means identifies any omissions or inconsistencies in the input data, the server uses the generation means to generate additional questions and notifies the user via the terminal.
[0766] For example, if a person receiving care gives a vague answer to the question, "Do you enjoy the food?", the server will consider their emotional state and suggest follow-up questions such as, "What kind of food do you particularly like?"
[0767] This system not only improves the accuracy of surveys but also reduces the psychological burden on users, supporting the smooth and reliable conduct of surveys.
[0768] The following describes the processing flow.
[0769] Step 1:
[0770] The device activates its voice acquisition function as soon as the on-site survey begins, recording the conversation clearly. Recording takes place in the background, and ambient noise is removed using the noise cancellation function.
[0771] Step 2:
[0772] The device transfers the acquired audio data to the speech recognition system in real time, where it is converted into text. The converted text is then saved in a format that can be processed immediately.
[0773] Step 3:
[0774] The server receives text data sent from the terminal and applies a summarization algorithm. This algorithm efficiently extracts important information and generates a summary.
[0775] Step 4:
[0776] The server automatically inputs the generated summary information into the questionnaire according to a standard format. Here, the information is accurately mapped to a pre-configured template.
[0777] Step 5:
[0778] Simultaneously, the device's emotion engine analyzes the conversation between the user and the person being cared for, recognizing their emotional state from their tone of voice and the flow of the conversation. Emotional data is accumulated in real time.
[0779] Step 6:
[0780] The server verifies the submitted questionnaires to check for any omissions or inconsistencies. During this process, it utilizes data from the emotion engine, focusing particularly on areas where emotional responses were strong.
[0781] Step 7:
[0782] If any omissions or inconsistencies are detected, the server will consider the emotional context and generate appropriate additional questions. The generated questions will be notified to the user via the terminal.
[0783] Step 8:
[0784] The user follows the instructions on the device, asks additional questions to the person being cared for, and records their responses. The newly obtained audio data is processed according to the aforementioned flow and reflected in the questionnaire.
[0785] Step 9:
[0786] Finally, the server verifies the content and securely stores the complete questionnaire in cloud storage or a database. This improves the accuracy of the survey records.
[0787] (Example 2)
[0788] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0789] The aim is to streamline the process of home visit surveys, from acquiring audio data to inputting information and analyzing emotional states, thereby ensuring accuracy and smooth progress of the surveys. Furthermore, the goal is to provide a system that reduces the psychological burden on care recipients and surveyors during the survey, enabling the acquisition of highly reliable data.
[0790] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0791] In this invention, the server includes an acquisition means for acquiring audio during a visit, a conversion means for converting audio into text information, a summarization means for extracting and summarizing important information, an emotion analysis means for evaluating emotional states through audio data analysis, and a generation means for verifying input data and generating additional inquiries. This enables efficient and accurate data acquisition during on-site surveys, as well as flexible survey progress that takes emotional states into consideration.
[0792] "Acquisition method" refers to a function for collecting audio data during visits while suppressing environmental noise.
[0793] The "conversion means" is a function that converts acquired speech into text information, and it uses a highly accurate language model.
[0794] A "summarization tool" is a function that extracts important information from text and converts it into a summarized format.
[0795] "Input method" refers to a function that automatically inputs summarized information into a standardized information format.
[0796] "Emotion analysis means" refers to a function that evaluates the emotional state of the user and the person they are talking to based on voice data.
[0797] A "detection mechanism" is a function that verifies input data, checks for omissions or inconsistencies, and adjusts the information.
[0798] A "generation mechanism" is a function that generates additional inquiries based on detected problems and missing information, and presents them to the interlocutor.
[0799] To implement this invention, a system is constructed that streamlines the on-site survey process and improves its accuracy by using the following hardware and software.
[0800] First, the device is equipped with an audio acquisition module and uses noise cancellation technology to acquire clear audio data. High-performance audio capture devices and noise reduction software are used for audio acquisition. For example, a microphone built into a smart device carried by the researcher and a signal processing algorithm that suppresses ambient noise could be considered.
[0801] Next, the device converts the audio data into text information using a conversion method. A cloud-based speech recognition API is used for speech recognition, and a multilingual language model handles dialects and intonation differences. For example, the Google Cloud Speech-to-Text API is used to convert audio data into text.
[0802] The server analyzes the received text information using a summarization mechanism, extracting and summarizing important information. Using natural language processing techniques, a specific algorithm automatically summarizes the key points of the information. For example, an algorithm that analyzes word frequency and relationships is used for information extraction.
[0803] On the other hand, the emotion analysis system analyzes voice data and evaluates the emotions of the user and the person they are speaking with. An emotion recognition algorithm is implemented to analyze the emotional state in real time based on the tone and content of the voice. This allows for the evaluation of changes in the emotions of the person being surveyed during the investigation.
[0804] The server uses detection methods to verify the input data and compare it to a template to check for any omissions or inconsistencies. By also considering the results of sentiment analysis, it is possible to adjust the questions if the user is experiencing stress.
[0805] Finally, the server generates additional queries to complement the detected problems using a generation mechanism. Using a generation AI model, it creates new questions based on appropriate prompts and presents them to the user via the terminal. An example of a prompt might be, "In a home visit, please describe how to analyze elderly people's feelings about food and identify their specific preferences."
[0806] This system improves the accuracy of on-site surveys, reduces the psychological burden on both the user and the interviewee, and enables smooth and reliable surveys.
[0807] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0808] Step 1:
[0809] The device activates its voice acquisition module at the start of the on-site survey and records voice data in real time while suppressing ambient noise using noise cancellation technology. The input is the conversation with the survey subject, and the output is a clear audio file. Specifically, the device's recording function is activated, and noise filtering is performed in the background.
[0810] Step 2:
[0811] The device converts recorded audio into text information using a conversion mechanism. The input is an audio file, and the output is text. Using a speech recognition API, a language model analyzes the audio waveform to convert it into text. Specifically, the audio is sent to a cloud service, and each audio segment is mapped to its corresponding text.
[0812] Step 3:
[0813] The server receives the converted text information, extracts important information using a summarization tool, and performs a summary. The input is text, and the output is summarized information. Natural language processing technology is used to analyze keywords and themes from the text and perform a summary. Specifically, a text analysis algorithm scans the content and selects sentences of high importance.
[0814] Step 4:
[0815] The server automatically inputs summarized information into a standardized information format. The input is summarized information, and the output is a standardized survey record. A database management system is used, and the summarization results are stored in the database according to the fields. Specifically, the information is automatically entered according to a predetermined format and registered as a survey record.
[0816] Step 5:
[0817] The emotion analysis tool analyzes the audio data and evaluates the emotional state of the user and the person they are speaking with. The input is the audio data, and the output is a score indicating the emotional state. The emotion identification algorithm analyzes the tone and pitch of the voice to identify changes in emotion. Specifically, the audio feature extraction tool scans the audio and assigns emotion labels.
[0818] Step 6:
[0819] The server uses detection mechanisms to check for missing or inconsistent data in the input and adjusts the data to take sentiment into account. The input consists of standardized survey records and sentiment scores, and the output is the corrected survey records. If inconsistencies are detected, the system re-evaluates the data and generates adjustment proposals. Specifically, the algorithm identifies the locations of missing data and applies corrective processing as needed.
[0820] Step 7:
[0821] The server uses a generation mechanism to generate additional inquiries regarding detected problems and notifies the user via the terminal. The input is inconsistent data and sentiment information, and the output is the additional inquiries. A generation AI model creates prompts and generates the necessary additional questions. For example, if there is an ambiguous answer, a question such as "What specific dishes do you like?" is automatically generated and suggested to the user via the terminal.
[0822] (Application Example 2)
[0823] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0824] In many physical stores, communication between customers and service staff significantly impacts the quality of service. However, it is difficult for service staff to accurately grasp the emotions and needs of all customers in real time, leading to inconsistencies in service quality. Furthermore, there are insufficient resources available to improve customer satisfaction by appropriately analyzing and addressing customer emotions and needs.
[0825] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0826] In this invention, the server includes a collection means for acquiring audio during a visit, a conversion means for converting the collected audio into text information, and an emotion analysis means for analyzing the audio and evaluating the emotional state. This makes it possible to improve the quality of communication between customers and service staff in physical stores, understand the customer's emotional state, and improve the response.
[0827] "Collection means" refers to a mechanism for acquiring voice data during visits.
[0828] "Conversion means" refers to technology for converting acquired audio data into text information.
[0829] A "summarization method" is a technique that extracts and summarizes important information from converted textual information.
[0830] "Input method" refers to technology that automatically inputs summarized information into a standard format.
[0831] A "detection means" is a mechanism for detecting omissions or inconsistencies in the input data.
[0832] "Generation means" refers to techniques for generating additional questions based on detected problems.
[0833] "Emotional analysis methods" refer to technologies that analyze voice to evaluate emotional states.
[0834] "Suggestion methods" refer to methods for improving customer service based on information obtained through emotion analysis.
[0835] The embodiment for carrying out the invention is configured as follows as a method for realizing a system to efficiently improve customer service in physical stores. This system has the function of collecting voice in real time, converting it into text information, and summarizing it. In addition, it analyzes the emotional state of customers from the voice data and uses that information to make suggestions for improving customer service.
[0836] The server uses speech recognition software to collect conversations between customers and service staff during visits. This collection method utilizes a recording device employing noise-canceling technology. Next, the collected audio data is converted into text information using natural language processing. At this stage, dialects and intonation are taken into consideration to generate accurate text data.
[0837] The converted textual information is summarized by a summarization algorithm, which extracts the most important information. This summarized information is then automatically entered according to a predetermined format. After input, the server uses detection means to check for any missing or inconsistent data, and generates additional questions as needed using generation means.
[0838] Furthermore, the server utilizes emotion analysis techniques to analyze text derived from the audio data and assess the customer's emotional state. This analysis employs a generative AI model that captures emotional changes based on voice tone and context.
[0839] For example, if a customer expresses positive feelings towards a product, the server can use this information to make further suggestions to the customer service representative, offering more helpful service. An example of a prompt would be, "Please provide sentiment analysis and summary of this conversation: 'I liked this product, but I have one more question.'"
[0840] This will improve the quality of customer service at physical stores and increase customer satisfaction.
[0841] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0842] Step 1:
[0843] The terminal activates the voice acquisition module and collects conversations between customers and service staff in real time during visits. Using noise cancellation technology, it acquires clear voice data while minimizing ambient noise. An audio signal is provided as input, and noise-reduced voice data is obtained as output.
[0844] Step 2:
[0845] The server processes the collected audio data using speech recognition software and converts it into text information. Speech recognition technology is used to convert audio information into text data. At this stage, the input is denoised audio data, and the output is recognized text data.
[0846] Step 3:
[0847] The server processes the converted character information using a summarization algorithm, extracting important information and creating a summary. This process uses natural language processing techniques to extract key points and shorten the text. The input is recognized character data, and the output is summarized text information.
[0848] Step 4:
[0849] The server automatically inputs summarized information into a pre-configured format. This format has a fixed data structure, and the data is automatically entered according to it. The input is summarized text information, and the output is a pre-filled datasheet according to the format.
[0850] Step 5:
[0851] The server analyzes the input data using detection methods to check for any missing or inconsistent data. The purpose of this process is to verify data integrity and identify any problems. The input is formatted data, and the output is the result of data verification.
[0852] Step 6:
[0853] Based on the problems detected by the server, the generation mechanism generates additional questions and notifies the user via the terminal. Here, questions are dynamically created if there are problems. The input is the verification result, and the output is a set of follow-up questions.
[0854] Step 7:
[0855] The server uses emotion analysis tools to evaluate the emotional state of text obtained from speech. A generative AI model is used to analyze voice tone and context. The input is textual information, and the output is emotional state data.
[0856] Step 8:
[0857] Based on the results of sentiment analysis, the server activates a suggestion mechanism to propose service improvements to the user. The results are returned as suggestions for the service to the user. The input is sentiment state data, and the output is improvement suggestions.
[0858] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0859] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0860] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0861] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0862] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0863] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0864] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0865] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0866] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0867] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0868] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0869] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0870] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0871] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0872] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0873] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0874] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0875] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0876] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0877] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0878] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0879] The following is further disclosed regarding the embodiments described above.
[0880] (Claim 1)
[0881] A means of acquiring audio during a visit,
[0882] A conversion means for converting acquired audio into text information,
[0883] A summarization means for summarizing the converted text information,
[0884] An input method that automatically inputs summarized information into a standard format,
[0885] A detection means for detecting missing or inconsistent input data,
[0886] A generation means for generating additional questions based on the detected problem,
[0887] A system that includes this.
[0888] (Claim 2)
[0889] The system according to claim 1, wherein the input means automatically inputs the summary information generated by the summarization means, associating it with each item of the questionnaire.
[0890] (Claim 3)
[0891] The system according to claim 1, wherein the detection means detects missing or inconsistent information by comparing the input data with a template.
[0892] "Example 1"
[0893] (Claim 1)
[0894] A device that acquires voice data during visits,
[0895] A device that converts acquired audio data into text information,
[0896] A device that summarizes converted text information and extracts important information,
[0897] A device that automatically inputs summarized information into a standardized format,
[0898] A device that detects missing or inconsistent information,
[0899] A device that generates questions to obtain additional information based on the detected problem,
[0900] A data processing device that includes a data processing device.
[0901] (Claim 2)
[0902] The data processing device according to claim 1, which automatically organizes and arranges the input information to correspond to items in a standard format.
[0903] (Claim 3)
[0904] The data processing device according to claim 1, wherein the detection device compares the input information with a standard format and identifies any missing or inconsistent information.
[0905] "Application Example 1"
[0906] (Claim 1)
[0907] A means of acquiring audio during a visit,
[0908] A conversion means for converting acquired audio into text information,
[0909] A summarization means for summarizing the converted text information,
[0910] An input method that automatically inputs summarized information into a standard format,
[0911] A detection means for detecting missing or inconsistent input data,
[0912] A recommendation method that acquires customer conversations and makes product recommendations,
[0913] A generation means for generating additional questions based on the detected problem,
[0914] A system that includes this.
[0915] (Claim 2)
[0916] The system according to claim 1, wherein the input means automatically inputs the summary information generated by the summarization means in correspondence with each item of the questionnaire, and the recommendation means makes product recommendations based on the conversation content.
[0917] (Claim 3)
[0918] The system according to claim 1, wherein the detection means compares input data with a template to detect any missing or inconsistent information, and the recommendation means performs a process of presenting the optimal product based on the extracted information.
[0919] "Example 2 of combining an emotion engine"
[0920] (Claim 1)
[0921] A means of acquiring audio during a visit,
[0922] A conversion means for converting acquired audio into text information,
[0923] A summarization means that summarizes the converted text information and extracts important information,
[0924] An input method for automatically entering summarized information into a standardized information table,
[0925] An emotion analysis means for evaluating the emotional state of the user and the person speaking to the other through voice data analysis,
[0926] A detection means that detects missing or inconsistent input data and adjusts the information based on emotional state,
[0927] A generation means that generates additional inquiries based on the detected problems and presents them to the interlocutor,
[0928] A system that includes this.
[0929] (Claim 2)
[0930] The system according to claim 1, wherein the input means automatically inputs the summary information generated by the summarization means, associating it with each item in the information table.
[0931] (Claim 3)
[0932] The system according to claim 1, wherein the detection means detects missing or inconsistent information by comparing the input data with a reference template.
[0933] "Application example 2 of combining emotional engines"
[0934] (Claim 1)
[0935] A collection method for acquiring audio during visits,
[0936] A conversion means for converting collected audio into text information,
[0937] A summarization means for summarizing the converted text information,
[0938] An input method that automatically inputs summarized information into a standard format,
[0939] A detection means for detecting missing or inconsistent input data,
[0940] A generation means for generating additional questions based on the detected problem,
[0941] An emotion analysis method that analyzes voice to evaluate emotional state,
[0942] A method for proposing improvements to customer service based on information from emotion analysis tools,
[0943] A system that includes this.
[0944] (Claim 2)
[0945] The system according to claim 1, wherein the input means automatically inputs the summary information generated by the summarization means, associating it with each item of the questionnaire.
[0946] (Claim 3)
[0947] The system according to claim 1, wherein the detection means detects missing or inconsistent information by comparing the input data with a template. [Explanation of Symbols]
[0948] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of acquiring audio during a visit, A conversion means for converting acquired audio into text information, A summarization means for summarizing the converted text information, An input method that automatically inputs summarized information into a standard format, A detection means for detecting missing or inconsistent input data, A generation means for generating additional questions based on the detected problem, A system that includes this.
2. The system according to claim 1, wherein the input means automatically inputs the summary information generated by the summarization means, associating it with each item of the questionnaire.
3. The system according to claim 1, wherein the detection means detects missing or inconsistent information by comparing the input data with a template.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A