system
A system converts and encrypts audio data to text, identifying personal information and securely transmitting it for analysis, addressing the risk of leakage and legal restrictions, enhancing data processing efficiency and customer satisfaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
The handling of personal information in voice data recorded during customer interactions poses a risk of leakage, especially when using external cloud services or AI systems, and legal restrictions on personal information protection hinder efficient and secure data processing.
A system that converts audio data to text using speech recognition, identifies personal information using natural language processing, encrypts it, and securely transmits it to external services while decrypting it only within the company for analysis.
Enables efficient data processing and analysis while preventing personal information leakage, allowing companies to comply with legal restrictions and enhance customer satisfaction.
Smart Images

Figure 2026070153000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern corporate activities, the handling of personal information included in voice data and the like that records interactions with customers is an important issue. In particular, when using external cloud services or AI systems, there is a risk of personal information leakage. For this reason, it is desired to process data efficiently and securely and promote digital transformation, but legal restrictions on personal information protection have become a major hurdle.
Means for Solving the Problems
[0005] This invention provides a system that receives audio data and converts it into text data using speech recognition technology. Furthermore, this system protects personal information by identifying it in the text data using natural language processing technology and encrypting it. It also has the ability to securely transmit the encrypted data to external cloud services or AI systems and decrypt it only within the company as needed. This enables efficient data processing using external services while preventing the leakage of personal information.
[0006] The "data receiving unit" is a configuration that has the function of acquiring audio data from an external source.
[0007] The "speech recognition unit" is a configuration that has the function of converting speech data into text data.
[0008] The "personal information detection unit" is a component that has the function of identifying personal information such as names and phone numbers from text data.
[0009] The "encryption processing unit" is a configuration that has an encryption function to convert personal information into a secure format.
[0010] The "data transmission unit" is a configuration that has the function of sending encrypted data to an external service.
[0011] "Natural language processing technology" refers to the technology of processing human language using computers, and in particular, it is a means of analyzing information in text.
[0012] The "decryption unit" is a component that has the function of restoring encrypted information to its original format. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the language used in the following description will be explained.
[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0019] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] The system of the present invention securely analyzes received audio data, minimizing the risk of personal information leakage while enabling integration with cloud services.
[0035] The system begins with the user recording calls and voice information and saving it as audio data on their device. This audio data is then sent to the server using a secure communication protocol.
[0036] The server stores the received audio data and converts it into text data using a speech recognition unit. This conversion makes the audio information searchable in text format, facilitating subsequent processing. Next, the server uses a personal information detection unit to identify personal information from the converted text data. Here, natural language processing technology is used to effectively extract important information such as names, addresses, phone numbers, and email addresses.
[0037] The identified personal information is encrypted by the server's encryption processing unit. This encryption protects the data when it is transmitted externally, preventing information leaks due to unauthorized access. The encrypted data is sent from the server to an external cloud service, where various analyses and processing are performed.
[0038] The results obtained from this external processing are returned to the server, which decrypts the data as needed and restores personal information. At this time, the decrypted information is kept only in a secure environment within the company, thus ensuring the protection of privacy.
[0039] As a concrete example, consider a case where a user records customer phone calls and uses a cloud service to analyze them. The server accurately transcribes the calls into text, masks customer personal information, and enables analysis of how customer interactions were conducted on the cloud. Based on the analysis results, specific improvement measures can then be fed back to the company.
[0040] This model allows companies to maximize the value of their data while complying with the law, enabling them to achieve both efficient operations and improved customer satisfaction.
[0041] The following describes the processing flow.
[0042] Step 1:
[0043] Users record phone calls or voice recordings and save the audio data to their devices. This audio data includes interactions and conversations with customers.
[0044] Step 2:
[0045] The terminal transmits voice data to the server using a secure communication protocol. Encryption is used during transmission to maintain data consistency and confidentiality.
[0046] Step 3:
[0047] The server stores the received audio data in its internal storage. When storing the data, it also stores metadata such as data identifiers and timestamps.
[0048] Step 4:
[0049] The server's speech recognition unit converts the received audio data into text data. In this process, a speech recognition algorithm is used to extract language from the audio file and convert it into the corresponding text.
[0050] Step 5:
[0051] The server's personal information detection unit analyzes the converted text data and uses internal natural language processing technology to identify personal information such as names, phone numbers, and addresses. The identified personal information is tagged and marked up for encryption.
[0052] Step 6:
[0053] The server's encryption processing unit encrypts the personal information marked up by the personal information detection unit. By using an encryption algorithm, the confidential information is converted into a format that cannot be identified.
[0054] Step 7:
[0055] The server's data transmission unit sends encrypted text data to an external cloud service. Since the data already has personal information protected, it is in a format suitable for external processing.
[0056] Step 8:
[0057] Cloud services receive encrypted data sent from servers and perform analysis and processing. For example, this can be used for evaluating customer service quality and analyzing trends.
[0058] Step 9:
[0059] The server receives processing results from the cloud service. The received result data is then fully utilized internally as needed.
[0060] Step 10:
[0061] The server's decryption unit decrypts encrypted data internally as needed. This process allows for the complete reconstruction of personal information only within the system, enabling detailed data analysis while protecting privacy.
[0062] (Example 1)
[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0064] There is a need for a method to securely digitize audio information while preventing the leakage of personal information, and to efficiently analyze and utilize the data results on external platforms. However, existing technologies make it difficult to balance information security and effective data utilization, and the risk of personal information leakage is a particular concern.
[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0066] In this invention, the server includes means for storing voice information with a data acquisition unit, means for converting the voice information into text information with a voice conversion unit, and means for converting the extracted important information into data with an information protection unit. This makes it possible to utilize advanced analysis on external platforms while ensuring the protection of personal information.
[0067] The term "data acquisition unit" refers to a device or function for collecting and storing audio information.
[0068] The term "voice conversion unit" refers to a device or function that converts voice information into digital data and then into text information.
[0069] The term "information detection unit" refers to a device or function used to identify and extract important data from textual information.
[0070] The "information protection unit" refers to the devices and functions that transform extracted sensitive data in order to securely protect it.
[0071] The "data transmission unit" refers to the device or function used to transmit the converted data to an external platform.
[0072] An "external platform" refers to an external computing environment that can analyze data received by the server and return the results.
[0073] The term "restoration unit" refers to a device or function that restores result data received from an external platform to its original state as needed.
[0074] "Generative algorithms" refer to methods used in data analysis to generate new insights and solutions.
[0075] This invention provides a system that processes voice information securely and efficiently, enabling analysis on external platforms while protecting personal information. The following describes embodiments for implementing this system.
[0076] First, the user records the call or audio information. The device saves this audio information in an appropriate digital format (e.g., WAV or MP3). When the device sends the saved audio information to the server, it uses a secure protocol such as TLS or SSL.
[0077] When the server receives audio information, it converts it into text information using a speech conversion unit. This conversion utilizes a speech recognition service (e.g., a general-purpose speech recognition API). The converted text information is then analyzed by an information detection unit to extract important data (e.g., name and address). Accuracy can be improved by utilizing natural language processing techniques at this stage.
[0078] The identified sensitive information is transformed by the Information Protection Unit into a secure, transmittable format. This information is then transmitted to an external platform by the Data Transmission Unit, where it is analyzed using generative algorithms. Through this analysis, users can gain new insights and improvement strategies.
[0079] As a concrete example, consider a case where a user records phone calls with customers and uses an external service to derive measures to improve customer satisfaction. The server transcribes the calls into text, analyzes them while protecting personal information, and then provides improvement suggestions based on this analysis.
[0080] An example of a prompt might be the instruction, "Analyze customer service calls and generate improvement suggestions." Based on this prompt, a generative AI model generates useful suggestions from the data. In this way, it becomes possible to maximize the use of corporate data while ensuring security, and to support efficient and effective decision-making.
[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0082] Step 1:
[0083] The user records phone calls and audio data. The input is the audio from the call. The device saves this audio in a digital format (e.g., a WAV file). The output is the saved digital audio file.
[0084] Step 2:
[0085] The terminal sends the stored audio file to the server. The input is a digital audio file. The terminal uses secure protocols such as TLS or SSL to ensure secure data transmission. The output is the audio file that has been securely delivered to the server.
[0086] Step 3:
[0087] The server processes the received audio file in its speech conversion unit and converts it into text information. The input is an audio data file. The server utilizes a speech recognition API to convert digital audio into text. The output is the converted text data.
[0088] Step 4:
[0089] The server analyzes textual information, and the information detection unit identifies personal information. The input is converted text data. Important information (e.g., name, address) is extracted using natural language processing technology. The output is the identified personal information.
[0090] Step 5:
[0091] The server encrypts and transforms the identified personal information using the information protection unit. The input is identified personal information. The server transforms the personal information using an encryption algorithm (e.g., AES). The output is the encrypted personal information.
[0092] Step 6:
[0093] The server transmits encrypted personal information to an external platform. The input is encrypted personal information. This data is securely transmitted to the external platform via the data transmission unit. The output is the data transmitted to the external platform.
[0094] Step 7:
[0095] The external platform analyzes data using a generative AI model and generates results. The input is encrypted personal information. The external system performs the analysis and generates new insights and suggestions. The output is the analysis results and suggested actions.
[0096] Step 8:
[0097] The server receives analysis results from an external platform and performs decryption. The input is the analysis results sent from the external platform. The server uses a recovery unit to decrypt the results as needed and return them to a usable format. The output is the decrypted analysis results.
[0098] Step 9:
[0099] Users or internal personnel take action to make decisions and implement improvements based on the decoded results. The input is the decoded analysis findings and suggestions. The output is the actions taken and the results obtained therefrom.
[0100] (Application Example 1)
[0101] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0102] In today's online commerce environment, streamlining customer service is a crucial challenge. However, analyzing customer call content carries the risk of personal information leaks, requiring appropriate protective measures. Furthermore, rapidly and accurately analyzing large amounts of call data and deriving concrete improvement measures is a technically challenging task.
[0103] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0104] In this invention, the server includes a data acquisition unit for acquiring voice information, a voice conversion unit for converting the voice information into text data, and an information detection unit for extracting personal information from the text data. This makes it possible to efficiently analyze customer call data while protecting personal information and use the results to improve customer service.
[0105] The "data acquisition unit" is a device that collects voice information and incorporates it as initial data.
[0106] The "voice conversion unit" is a device that converts collected voice information into text data.
[0107] The "information detection unit" is a device that extracts personal information from text data and identifies the necessary information.
[0108] The "processing unit" is a device that encrypts the extracted personal information to ensure data security.
[0109] The "transmission unit" is a device used to transmit encrypted information to an external system.
[0110] The "analysis unit" is a device that uses analysis results obtained from external systems to improve internal operational efficiency.
[0111] The "decryption unit" is a device that restores encrypted personal information and reconstructs the original information in a secure environment.
[0112] This invention realizes a system that efficiently acquires and analyzes voice information and utilizes the results. First, the user makes a call with a customer through smart glasses. Voice information is collected by the data acquisition unit. The smart glasses can be wearable devices available on the market, such as Google® Glass®.
[0113] The collected audio information is converted into text data through a speech-to-text unit. This process can utilize speech recognition technologies such as the Google Speech-to-Text API. The converted text data is then processed by an information detection unit, which uses natural language processing techniques to identify personal information. NLP tools such as SpaCy are useful in this step.
[0114] Personal information identified by the information detection unit is encrypted by the processing unit to ensure security. Encryption libraries such as OpenSSL can be used for this encryption. The encrypted information is then transmitted to an analysis service on the cloud via the transmission unit.
[0115] On the cloud side, data is analyzed, and the results are fed back internally through the analysis unit. For example, areas for improvement in customer service are identified. The analysis unit can use a generative AI model to generate prompt messages and suggest improvements.
[0116] As a concrete example, consider a case where a customer inquires about "order cancellation" at a support center. This inquiry is recorded as audio, converted to text, and then keywords such as "order" and "cancellation" are extracted by the information detection unit. As a result, if similar inquiries occur frequently, common patterns in the reasons for cancellation can be identified, which can be used to improve support operations.
[0117] An example of a prompt for the generating AI model could be text such as, "Analyze the patterns in why customers cancel orders." This allows the system to efficiently analyze data and generate specific improvement suggestions.
[0118] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0119] Step 1:
[0120] The user makes phone calls with customers using smart glasses. The calls are collected as audio information by the data acquisition unit. The audio information is stored as raw data in the smart glasses' memory.
[0121] Step 2:
[0122] The terminal sends the collected audio information to the server. The server uses a speech conversion unit to convert the audio information into text data. This process utilizes speech recognition software (e.g., Google Speech-to-Text API). The audio data (input) is converted into text data (output).
[0123] Step 3:
[0124] The server processes the text data in its information detection unit. At this stage, natural language processing technology (e.g., SpaCy) is used to identify personal information. Personal information such as names and addresses (output) is extracted from the text data (input).
[0125] Step 4:
[0126] The server encrypts the identified personal information in its processing unit. Using an encryption library (e.g., OpenSSL), the personal information (input) is converted into secure encrypted data (output).
[0127] Step 5:
[0128] The server sends encrypted data to the cloud analysis service via the transmission unit. The cloud system performs the analysis and generates prompt messages using an AI model that generates the results. The encrypted data (input) is converted into analysis results (output), and prompt messages are generated.
[0129] Step 6:
[0130] The analysis results returned from the cloud are fed back by the analysis unit on the server. The server then identifies specific areas for improvement to enhance internal operational efficiency. The analysis results (input) are then used as actual improvement suggestions (output).
[0131] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0132] The system in this invention has advanced data processing capabilities that simultaneously evaluate the user's emotional state in addition to analyzing voice data. This invention aims to improve communication with customers by ensuring the secure handling of personal information and enabling data analysis that takes user emotions into consideration.
[0133] The system operation begins with the user recording audio data. The user uses a smartphone or dedicated device to record phone calls and conversations, saving the data to the device. The recorded audio data is then transmitted from the device to the server using a secure protocol.
[0134] The server stores the received audio data in a database and simultaneously converts it into text data using a speech recognition unit. This conversion makes the audio information machine-processable. When analyzing the text data, the server utilizes a personal information detection unit to identify personal information such as names and addresses, and securely protects this information using an encryption processing unit.
[0135] Furthermore, this system is equipped with an emotion engine that estimates the user's emotions from voice data. The emotion engine analyzes the tone, speed, and rhythm of the voice and classifies the user's emotional state into several categories (e.g., joy, sadness, anger). This information is stored in a database as an essential element for improving customer service.
[0136] Encrypted data and analysis results from the emotion engine are securely transmitted to an external cloud service. Personal information is masked during this process, enabling in-depth analysis by the external service. Once the external processing results are returned to the server, the system decrypts the data as needed and uses that information to develop detailed customer service improvement strategies internally.
[0137] For example, when a recording of a customer's phone call to customer support is analyzed, the customer's emotions can be identified as "dissatisfaction" from the recording, and the analysis results are provided while appropriately protecting the personal information of the call. This allows the support team to take concrete approaches to improve the quality of customer service.
[0138] This model allows companies to protect personal information while providing sophisticated services that take customer emotions into consideration, thereby securing a competitive advantage in the market.
[0139] The following describes the processing flow.
[0140] Step 1:
[0141] The user records audio data. Using a smartphone or dedicated device, calls and conversations are recorded in real time.
[0142] Step 2:
[0143] The device transmits the recorded audio data to the server via a secure protocol. The data is encrypted during transmission to maintain confidentiality.
[0144] Step 3:
[0145] The server stores the received audio data in a database. The data is also accompanied by metadata such as an identifier and the date and time of recording.
[0146] Step 4:
[0147] The server's speech recognition unit converts the audio data into text data. The converted text becomes the basic data for analyzing the conversation content as written text.
[0148] Step 5:
[0149] The server's personal information detection unit analyzes text data and uses natural language processing technology to identify personal information such as names and addresses. The identified information is then marked for encryption.
[0150] Step 6:
[0151] The server's encryption processing unit encrypts the identified personal information. The encrypted personal information is rendered harmless, ensuring security when exchanging data with external parties.
[0152] Step 7:
[0153] The server's emotion engine analyzes the voice data and estimates the user's emotional state from the tone and rhythm of their voice. The estimated emotion is then classified into categories such as "joy" or "dissatisfaction."
[0154] Step 8:
[0155] The server sends encrypted text data and sentiment analysis results to an external cloud service. This external service analyzes the data and generates detailed insights.
[0156] Step 9:
[0157] The server receives analysis results returned from an external service. These results may include, for example, customer sentiment trends and areas for improvement in service delivery.
[0158] Step 10:
[0159] The server's decryption unit decrypts the data as needed. This is used for internal verification and specific customer support. This ensures that the information can be used while maintaining security.
[0160] (Example 2)
[0161] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0162] In analyzing voice data, it is necessary to accurately estimate the user's emotional state, securely protect personal information, and quickly and efficiently integrate with external services. Conventional systems have challenges in the accuracy of emotion estimation and the protection of personal information, and rapid external integration is required.
[0163] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0164] In this invention, the server includes a data acquisition unit for acquiring voice information, a voice conversion unit for converting the voice information into text information, a personal information identification unit for identifying personal information from the text information, an emotion analysis unit for estimating emotional states from voice information, an information protection unit for encrypting identified personal information, and a data communication unit for transmitting the encrypted information and emotional states to an external service. This enables data analysis that takes into account the user's emotional state, secure protection of personal information, and rapid external collaboration.
[0165] A "data acquisition unit" is a device or function that collects voice information and incorporates it into the system.
[0166] A "speech conversion unit" refers to a device or technology that has the function of analyzing acquired speech information and converting it into text information.
[0167] A "personal information identification unit" is a device or process that has the function of accurately extracting and identifying information related to a specific individual from textual information.
[0168] "Information protection" refers to a device or technology that has the function of encrypting identified personal information in order to securely protect that information.
[0169] A "sentiment analysis unit" refers to a device or software that has the function of estimating and classifying the emotional state of a speaker by analyzing audio information.
[0170] A "data communication unit" refers to a device or function equipped with communication means for transmitting processed information to an external service.
[0171] The system based on this invention analyzes voice information through collaboration between the user, terminal, and server, securely protects personal information, evaluates the user's emotional state, and integrates with external services. The system aims to improve service quality and facilitate smooth communication with customers by ensuring that each department operates appropriately.
[0172] First, the user records audio via their smartphone or a dedicated device. The recorded audio data is then sent to the server by the device using a secure protocol. During this process, the device encrypts the audio data to ensure its security.
[0173] When the server receives data, it first stores it in a database and then uses a speech conversion unit (e.g., speech recognition software) to convert the speech data into text information. Next, it uses a personal information identification unit to identify personal information from the text information. At this stage, natural language processing technology is utilized, and the identified personal information is encrypted by the information protection unit.
[0174] Next, the emotion analysis unit installed on the server analyzes the voice data and infers the emotional state. The emotion engine classifies the information into categories such as "joy" and "sadness" based on the tone and speed of the voice. The results are stored in a database and used for marketing and customer service.
[0175] Furthermore, the server securely transmits encrypted personal information and sentiment analysis data to an external cloud service via the data communication unit. This enables detailed analysis by the external service, and the analysis results are sent back to the server for internal use.
[0176] As a concrete example, consider a scenario where a customer calls customer support and the conversation is recorded. In this situation, the system recognizes the emotion of the call as "dissatisfaction" and provides analysis results while protecting the personal information of the caller. In this way, the support team can develop specific measures to improve the quality of customer service.
[0177] Examples of prompt statements in generative AI models include the following:
[0178] "We estimate customer emotions based on voice data and propose specific countermeasures."
[0179] Such systems enable companies to provide sophisticated services that take customer emotions into consideration while strictly protecting personal information, thereby maintaining a competitive advantage in the market.
[0180] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0181] Step 1:
[0182] The user records audio using a smartphone or dedicated device. The input is the user's voice, and the output is an audio data file. The microphone on the device captures the sound and saves it in a file format.
[0183] Step 2:
[0184] The terminal sends recorded audio data to the server. The input is an audio data file, and the output is data transfer to the server. The terminal encrypts the data using a secure protocol and sends it to the specified server address.
[0185] Step 3:
[0186] The server saves the received audio data to a database. The input is the audio data transferred from the terminal, and the output is a file stored in the database. The server verifies the integrity of the data before storing it in the database.
[0187] Step 4:
[0188] The server converts audio data into text information using a speech conversion unit. The input is an audio data file, and the output is text data. Speech recognition software analyzes the audio into text and generates text information.
[0189] Step 5:
[0190] The server uses a personal information identification unit to identify personal information from text data. The input is text data, and the output is the detected personal information. Natural language processing technology is used to extract names, addresses, and other information.
[0191] Step 6:
[0192] The server uses the information protection unit to encrypt the identified personal information. The input is the detected personal information, and the output is the encrypted personal information. An encryption algorithm is applied to ensure security.
[0193] Step 7:
[0194] The server uses an emotion analysis unit to estimate emotional states from voice data. The input is voice data, and the output is a classified emotional state. It analyzes the tone and speed of the voice to estimate emotions such as "joy" or "sadness."
[0195] Step 8:
[0196] The server transmits encrypted information and emotional states to an external service via a data communication unit. The input is encrypted personal information and emotional states, and the output is data transfer to an external service. Data is securely shared using an API.
[0197] (Application Example 2)
[0198] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0199] In today's world, there is a need to accurately monitor emotional changes while ensuring individual safety. In particular, there is a need for a system that can proactively detect and warn of risks associated with rapid emotional shifts. However, existing voice analysis technologies struggle to evaluate a user's emotional state in real time and generate appropriate alerts while maintaining the protection of personal information. Technologies are needed to overcome these limitations.
[0200] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0201] In this invention, the server includes means for acquiring acoustic information, means for converting acoustic information into text information, and means for analyzing the acoustic information to evaluate the user's emotional state. This makes it possible to detect emotional states from acoustic information and generate warnings for abnormal emotional changes while protecting personal information.
[0202] The "data acquisition unit" is a component that has the function of collecting acoustic information from the user's environment.
[0203] The "speech recognition unit" is a component that appropriately converts acquired acoustic information into text information and makes it into an analyzable format.
[0204] The "personal information detection unit" is the part of the system that has the function of identifying information that can identify an individual from textual information and analyzing it securely.
[0205] The "encryption processing unit" is the part of the system that has the function of encrypting identified personal information to protect it from unauthorized access by others.
[0206] The "emotion analysis unit" is a dedicated component that analyzes the tone and rhythm of the user's voice from acoustic information and evaluates their emotional state.
[0207] The "data transmission unit" is a configuration that has the function of transmitting processed data and encrypted information to external functions or services.
[0208] The "warning generation unit" is a component that has the function of issuing an appropriate warning when the user's emotional state exceeds a predetermined threshold.
[0209] To implement this invention, a system is needed that has a series of processes for collecting acoustic information, analyzing that information, and evaluating the user's emotional state. First, acoustic information is collected from the user's surroundings using a terminal such as a smartphone or a dedicated device. This data is securely transmitted to a server through a data acquisition unit. On the server, speech recognition software (e.g., Google Speech-to-Text API) is used to convert the acoustic information into text information.
[0210] The server analyzes textual information using natural language processing technology, and the personal information identified by the personal information detection unit is protected by the encryption processing unit. Meanwhile, the emotion analysis unit uses a generative AI model (e.g., a custom model using TENSORFLOW®) to evaluate the user's emotional state from acoustic information. This process includes the analysis of sound tone and rhythm, and emotions are classified into multiple categories.
[0211] If the emotional state exceeds a set threshold, the server sends an alert to emergency contacts via the warning generation unit. This procedure enables real-time emotion monitoring and warning generation based on acoustic information while maintaining the security of personal information.
[0212] For example, if a user is experiencing significant stress at home, the system immediately detects this change in emotion and sends a warning message stating, "The user may be experiencing stress." This warning is then sent via a smartphone app to pre-configured emergency contacts.
[0213] An example of a prompt message would be: "Use this prompt message to monitor the user's emotional changes. Analyze the emotional information obtained from the voice data and issue appropriate actions according to the urgency." This allows the system to effectively track the user's emotional state and take necessary actions.
[0214] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0215] Step 1:
[0216] The terminal collects acoustic information from the user's surroundings using its built-in microphone. This acoustic information becomes the input data. The terminal then prepares to transmit this data to the server via the data acquisition unit.
[0217] Step 2:
[0218] The server receives the acoustic information transmitted from the terminal in its data acquisition unit. Next, it uses speech recognition software (e.g., Google Speech-to-Text API) to convert this acoustic information into text information. This conversion results in the text information becoming the output data.
[0219] Step 3:
[0220] The server analyzes the text information converted by the speech recognition unit using natural language processing technology, and extracts personally identifiable information in the personal information detection unit. This personal information is the output of the processing and is then securely protected by the encryption processing unit.
[0221] Step 4:
[0222] The server uses an emotion analysis unit to analyze the tone and rhythm of the voice from the acoustic information, and utilizes a generative AI model to evaluate the user's emotional state. In this process, the acoustic information is the input data, and the emotion evaluation result is obtained as output data.
[0223] Step 5:
[0224] The server determines whether the emotional state has exceeded a set threshold based on the emotion evaluation results obtained from the emotion analysis unit. If the threshold is exceeded, the warning generation unit sends an alert to the emergency contact. At this time, the warning alert is output using the emotion evaluation results as input.
[0225] Step 6:
[0226] The user receives an alert from the server and takes the appropriate action according to the instructions in the prompt. In this step, the action based on the prompt appears as output.
[0227] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0228] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0229] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0230] [Second Embodiment]
[0231] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0232] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0233] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0234] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0235] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0236] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0237] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0238] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0239] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0240] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0241] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0242] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0243] The system of the present invention securely analyzes received audio data, minimizing the risk of personal information leakage while enabling integration with cloud services.
[0244] The system begins with the user recording calls and voice information and saving it as audio data on their device. This audio data is then sent to the server using a secure communication protocol.
[0245] The server stores the received audio data and converts it into text data using a speech recognition unit. This conversion makes the audio information searchable in text format, facilitating subsequent processing. Next, the server uses a personal information detection unit to identify personal information from the converted text data. Here, natural language processing technology is used to effectively extract important information such as names, addresses, phone numbers, and email addresses.
[0246] The identified personal information is encrypted by the server's encryption processing unit. This encryption protects the data when it is transmitted externally, preventing information leaks due to unauthorized access. The encrypted data is sent from the server to an external cloud service, where various analyses and processing are performed.
[0247] The results obtained from this external processing are returned to the server, which decrypts the data as needed and restores personal information. At this time, the decrypted information is kept only in a secure environment within the company, thus ensuring the protection of privacy.
[0248] As a concrete example, consider a case where a user records customer phone calls and uses a cloud service to analyze them. The server accurately transcribes the calls into text, masks customer personal information, and enables analysis of how customer interactions were conducted on the cloud. Based on the analysis results, specific improvement measures can then be fed back to the company.
[0249] This model allows companies to maximize the value of their data while complying with the law, enabling them to achieve both efficient operations and improved customer satisfaction.
[0250] The following describes the processing flow.
[0251] Step 1:
[0252] Users record phone calls or voice recordings and save the audio data to their devices. This audio data includes interactions and conversations with customers.
[0253] Step 2:
[0254] The terminal transmits voice data to the server using a secure communication protocol. Encryption is used during transmission to maintain data consistency and confidentiality.
[0255] Step 3:
[0256] The server stores the received audio data in its internal storage. When storing the data, it also stores metadata such as data identifiers and timestamps.
[0257] Step 4:
[0258] The server's speech recognition unit converts the received audio data into text data. In this process, a speech recognition algorithm is used to extract language from the audio file and convert it into the corresponding text.
[0259] Step 5:
[0260] The server's personal information detection unit analyzes the converted text data and uses internal natural language processing technology to identify personal information such as names, phone numbers, and addresses. The identified personal information is tagged and marked up for encryption.
[0261] Step 6:
[0262] The server's encryption processing unit encrypts the personal information marked up by the personal information detection unit. By using an encryption algorithm, the confidential information is converted into a format that cannot be identified.
[0263] Step 7:
[0264] The server's data transmission unit sends encrypted text data to an external cloud service. Since the data already has personal information protected, it is in a format suitable for external processing.
[0265] Step 8:
[0266] Cloud services receive encrypted data sent from servers and perform analysis and processing. For example, this can be used for evaluating customer service quality and analyzing trends.
[0267] Step 9:
[0268] The server receives processing results from the cloud service. The received result data is then fully utilized internally as needed.
[0269] Step 10:
[0270] The server's decryption unit decrypts encrypted data internally as needed. This process allows for the complete reconstruction of personal information only within the system, enabling detailed data analysis while protecting privacy.
[0271] (Example 1)
[0272] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0273] There is a need for a method to securely digitize audio information while preventing the leakage of personal information, and to efficiently analyze and utilize the data results on external platforms. However, existing technologies make it difficult to balance information security and effective data utilization, and the risk of personal information leakage is a particular concern.
[0274] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0275] In this invention, the server includes a data acquisition unit for storing voice information, a voice conversion unit for converting the voice information into character information, and an information protection unit for converting the extracted important information. This enables the utilization of advanced analysis on an external platform while ensuring the protection of personal information.
[0276] The "data acquisition unit" refers to a device or function for collecting and storing voice information.
[0277] The "voice conversion unit" refers to a device or function for converting voice information into digital data and then into character information.
[0278] The "information detection unit" refers to a device or function for identifying and extracting important data from character information.
[0279] The "information protection unit" refers to a device or function for converting data to securely protect the extracted important data.
[0280] The "data transmission unit" refers to a device or function for transmitting the converted data to an external platform.
[0281] The "external platform" refers to an external computing environment that can analyze the data received on the server side and return the results.
[0282] The "restoration unit" refers to a device or function for restoring the result data received from an external platform to its original state as needed.
[0283] The "generative algorithm" refers to a method for generating new insights and solutions in data analysis.
[0284] This invention is a system that safely and efficiently processes voice information, enables analysis on an external platform while protecting personal information. The following shows the embodiments for implementing this system.
[0285] First, the user records a call or voice information. The terminal saves this voice information in an appropriate digital format (e.g., WAV or MP3). When the terminal transmits the saved voice information to the server, it uses a secure protocol such as TLS or SSL.
[0286] When the server receives the voice information, it converts it into character information using a voice conversion unit. For this conversion, a voice recognition service (e.g., a general-purpose voice recognition API) is used. The converted character information is then analyzed by an information detection unit, and important data (e.g., name or address) is extracted. Here, the accuracy is improved by leveraging natural language processing technology.
[0287] The identified important information is data-converted by the information protection unit and converted into a form that can be safely transmitted. This information is transmitted to an external platform by a data transmission unit, and analysis using a generative algorithm is performed on the external platform. Through this analysis, the user can obtain new insights and improvement measures.
[0288] As a specific example, consider the case where a user records a call with a customer and uses an external service to derive measures to improve customer satisfaction. The server converts the call into text, performs analysis while protecting personal information, and improvement proposals are made based on this analysis.
[0289] As an example of a prompt sentence, an instruction such as "Analyze the customer service call and create improvement proposals" can be considered. Based on this prompt, the generative AI model generates useful proposals from the data. In this way, it is possible to maximize the utilization of enterprise data while ensuring security and support efficient and effective decision-making.
[0290] The flow of the specific process in Example 1 will be described using FIG. 11.
[0291] Step 1:
[0292] The user records phone calls and audio data. The input is the audio from the call. The device saves this audio in a digital format (e.g., a WAV file). The output is the saved digital audio file.
[0293] Step 2:
[0294] The terminal sends the stored audio file to the server. The input is a digital audio file. The terminal uses secure protocols such as TLS or SSL to ensure secure data transmission. The output is the audio file that has been securely delivered to the server.
[0295] Step 3:
[0296] The server processes the received audio file in its speech conversion unit and converts it into text information. The input is an audio data file. The server utilizes a speech recognition API to convert digital audio into text. The output is the converted text data.
[0297] Step 4:
[0298] The server analyzes textual information, and the information detection unit identifies personal information. The input is converted text data. Important information (e.g., name, address) is extracted using natural language processing technology. The output is the identified personal information.
[0299] Step 5:
[0300] The server encrypts and transforms the identified personal information using the information protection unit. The input is identified personal information. The server transforms the personal information using an encryption algorithm (e.g., AES). The output is the encrypted personal information.
[0301] Step 6:
[0302] The server sends the encrypted personal information to an external platform. The input is the encrypted personal information. Through the data transmission unit, this data is securely sent to the external platform. The output is the data sent to the external platform.
[0303] Step 7:
[0304] The external platform analyzes the data using a generative AI model and generates results. The input is the encrypted personal information. The external system performs the analysis and generates new insights and suggestions. The output is the analysis results and the proposed actions.
[0305] Step 8:
[0306] The server receives the analysis results from the external platform and decrypts them. The input is the analysis results sent from the external platform. The server uses the restoration unit to decrypt the results as needed and returns them to a usable form. The output is the decrypted analysis results.
[0307] Step 9:
[0308] The user or the in-house responsible person takes actions for decision-making and implementing improvement measures based on the decrypted results. The input is the decrypted analysis findings and suggestions. The output is the actions taken and the resulting achievements.
[0309] (Application Example 1)
[0310] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0311] In today's online commerce environment, streamlining customer service is a crucial challenge. However, analyzing customer call content carries the risk of personal information leaks, requiring appropriate protective measures. Furthermore, rapidly and accurately analyzing large amounts of call data and deriving concrete improvement measures is a technically challenging task.
[0312] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0313] In this invention, the server includes a data acquisition unit for acquiring voice information, a voice conversion unit for converting the voice information into text data, and an information detection unit for extracting personal information from the text data. This makes it possible to efficiently analyze customer call data while protecting personal information and use the results to improve customer service.
[0314] The "data acquisition unit" is a device that collects voice information and incorporates it as initial data.
[0315] The "voice conversion unit" is a device that converts collected voice information into text data.
[0316] The "information detection unit" is a device that extracts personal information from text data and identifies the necessary information.
[0317] The "processing unit" is a device that encrypts the extracted personal information to ensure data security.
[0318] The "transmission unit" is a device used to transmit encrypted information to an external system.
[0319] The "analysis unit" is a device that uses analysis results obtained from external systems to improve internal operational efficiency.
[0320] The "decryption unit" is a device that restores encrypted personal information and reconstructs the original information in a secure environment.
[0321] This invention realizes a system that efficiently acquires and analyzes voice information and utilizes the results. First, the user makes a call with a customer through smart glasses. Voice information is collected by the data acquisition unit. The smart glasses can be wearable devices available on the market, such as Google Glass.
[0322] The collected audio information is converted into text data through a speech-to-text unit. This process can utilize speech recognition technologies such as the Google Speech-to-Text API. The converted text data is then processed by an information detection unit, which uses natural language processing techniques to identify personal information. NLP tools such as SpaCy are useful in this step.
[0323] Personal information identified by the information detection unit is encrypted by the processing unit to ensure security. Encryption libraries such as OpenSSL can be used for this encryption. The encrypted information is then transmitted to an analysis service on the cloud via the transmission unit.
[0324] On the cloud side, data is analyzed, and the results are fed back internally through the analysis unit. For example, areas for improvement in customer service are identified. The analysis unit can use a generative AI model to generate prompt messages and suggest improvements.
[0325] As a concrete example, consider a case where a customer inquires about "order cancellation" at a support center. This inquiry is recorded as audio, converted to text, and then keywords such as "order" and "cancellation" are extracted by the information detection unit. As a result, if similar inquiries occur frequently, common patterns in the reasons for cancellation can be identified, which can be used to improve support operations.
[0326] An example of a prompt for the generating AI model could be text such as, "Analyze the patterns in why customers cancel orders." This allows the system to efficiently analyze data and generate specific improvement suggestions.
[0327] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0328] Step 1:
[0329] The user makes phone calls with customers using smart glasses. The calls are collected as audio information by the data acquisition unit. The audio information is stored as raw data in the smart glasses' memory.
[0330] Step 2:
[0331] The terminal sends the collected audio information to the server. The server uses a speech conversion unit to convert the audio information into text data. This process utilizes speech recognition software (e.g., Google Speech-to-Text API). The audio data (input) is converted into text data (output).
[0332] Step 3:
[0333] The server processes the text data in its information detection unit. At this stage, natural language processing technology (e.g., SpaCy) is used to identify personal information. Personal information such as names and addresses (output) is extracted from the text data (input).
[0334] Step 4:
[0335] The server encrypts the identified personal information in its processing unit. Using an encryption library (e.g., OpenSSL), the personal information (input) is converted into secure encrypted data (output).
[0336] Step 5:
[0337] The server sends encrypted data to the cloud analysis service via the transmission unit. The cloud system performs the analysis and generates prompt messages using an AI model that generates the results. The encrypted data (input) is converted into analysis results (output), and prompt messages are generated.
[0338] Step 6:
[0339] The analysis results returned from the cloud are fed back by the analysis unit on the server. The server then identifies specific areas for improvement to enhance internal operational efficiency. The analysis results (input) are then used as actual improvement suggestions (output).
[0340] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0341] The system in this invention has advanced data processing capabilities that simultaneously evaluate the user's emotional state in addition to analyzing voice data. This invention aims to improve communication with customers by ensuring the secure handling of personal information and enabling data analysis that takes user emotions into consideration.
[0342] The system operation begins with the user recording audio data. The user uses a smartphone or dedicated device to record phone calls and conversations, saving the data to the device. The recorded audio data is then transmitted from the device to the server using a secure protocol.
[0343] The server stores the received audio data in a database and simultaneously converts it into text data using a speech recognition unit. This conversion makes the audio information machine-processable. When analyzing the text data, the server utilizes a personal information detection unit to identify personal information such as names and addresses, and securely protects this information using an encryption processing unit.
[0344] Furthermore, this system is equipped with an emotion engine that estimates the user's emotions from voice data. The emotion engine analyzes the tone, speed, and rhythm of the voice and classifies the user's emotional state into several categories (e.g., joy, sadness, anger). This information is stored in a database as an essential element for improving customer service.
[0345] Encrypted data and analysis results from the emotion engine are securely transmitted to an external cloud service. Personal information is masked during this process, enabling in-depth analysis by the external service. Once the external processing results are returned to the server, the system decrypts the data as needed and uses that information to develop detailed customer service improvement strategies internally.
[0346] For example, when a recording of a customer's phone call to customer support is analyzed, the customer's emotions can be identified as "dissatisfaction" from the recording, and the analysis results are provided while appropriately protecting the personal information of the call. This allows the support team to take concrete approaches to improve the quality of customer service.
[0347] This model allows companies to protect personal information while providing sophisticated services that take customer emotions into consideration, thereby securing a competitive advantage in the market.
[0348] The following describes the processing flow.
[0349] Step 1:
[0350] The user records audio data. Using a smartphone or dedicated device, calls and conversations are recorded in real time.
[0351] Step 2:
[0352] The device transmits the recorded audio data to the server via a secure protocol. The data is encrypted during transmission to maintain confidentiality.
[0353] Step 3:
[0354] The server stores the received audio data in a database. The data is also accompanied by metadata such as an identifier and the date and time of recording.
[0355] Step 4:
[0356] The server's speech recognition unit converts the audio data into text data. The converted text becomes the basic data for analyzing the conversation content as written text.
[0357] Step 5:
[0358] The server's personal information detection unit analyzes text data and uses natural language processing technology to identify personal information such as names and addresses. The identified information is then marked for encryption.
[0359] Step 6:
[0360] The server's encryption processing unit encrypts the identified personal information. The encrypted personal information is rendered harmless, ensuring security when exchanging data with external parties.
[0361] Step 7:
[0362] The server's emotion engine analyzes the voice data and estimates the user's emotional state from the tone and rhythm of their voice. The estimated emotion is then classified into categories such as "joy" or "dissatisfaction."
[0363] Step 8:
[0364] The server sends encrypted text data and sentiment analysis results to an external cloud service. This external service analyzes the data and generates detailed insights.
[0365] Step 9:
[0366] The server receives analysis results returned from an external service. These results may include, for example, customer sentiment trends and areas for improvement in service delivery.
[0367] Step 10:
[0368] The server's decryption unit decrypts the data as needed. This is used for internal verification and specific customer support. This ensures that the information can be used while maintaining security.
[0369] (Example 2)
[0370] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0371] In analyzing voice data, it is necessary to accurately estimate the user's emotional state, securely protect personal information, and quickly and efficiently integrate with external services. Conventional systems have challenges in the accuracy of emotion estimation and the protection of personal information, and rapid external integration is required.
[0372] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0373] In this invention, the server includes a data acquisition unit for acquiring voice information, a voice conversion unit for converting the voice information into text information, a personal information identification unit for identifying personal information from the text information, an emotion analysis unit for estimating emotional states from voice information, an information protection unit for encrypting identified personal information, and a data communication unit for transmitting the encrypted information and emotional states to an external service. This enables data analysis that takes into account the user's emotional state, secure protection of personal information, and rapid external collaboration.
[0374] A "data acquisition unit" is a device or function that collects voice information and incorporates it into the system.
[0375] A "speech conversion unit" refers to a device or technology that has the function of analyzing acquired speech information and converting it into text information.
[0376] A "personal information identification unit" is a device or process that has the function of accurately extracting and identifying information related to a specific individual from textual information.
[0377] "Information protection" refers to a device or technology that has the function of encrypting identified personal information in order to securely protect that information.
[0378] A "sentiment analysis unit" refers to a device or software that has the function of estimating and classifying the emotional state of a speaker by analyzing audio information.
[0379] A "data communication unit" refers to a device or function equipped with communication means for transmitting processed information to an external service.
[0380] The system based on this invention analyzes voice information through collaboration between the user, terminal, and server, securely protects personal information, evaluates the user's emotional state, and integrates with external services. The system aims to improve service quality and facilitate smooth communication with customers by ensuring that each department operates appropriately.
[0381] First, the user records audio via their smartphone or a dedicated device. The recorded audio data is then sent to the server by the device using a secure protocol. During this process, the device encrypts the audio data to ensure its security.
[0382] When the server receives data, it first stores it in a database and then uses a speech conversion unit (e.g., speech recognition software) to convert the speech data into text information. Next, it uses a personal information identification unit to identify personal information from the text information. At this stage, natural language processing technology is utilized, and the identified personal information is encrypted by the information protection unit.
[0383] Next, the emotion analysis unit installed on the server analyzes the voice data and infers the emotional state. The emotion engine classifies the information into categories such as "joy" and "sadness" based on the tone and speed of the voice. The results are stored in a database and used for marketing and customer service.
[0384] Furthermore, the server securely transmits encrypted personal information and sentiment analysis data to an external cloud service via the data communication unit. This enables detailed analysis by the external service, and the analysis results are sent back to the server for internal use.
[0385] As a concrete example, consider a scenario where a customer calls customer support and the conversation is recorded. In this situation, the system recognizes the emotion of the call as "dissatisfaction" and provides analysis results while protecting the personal information of the caller. In this way, the support team can develop specific measures to improve the quality of customer service.
[0386] Examples of prompt statements in generative AI models include the following:
[0387] "We estimate customer emotions based on voice data and propose specific countermeasures."
[0388] Such systems enable companies to provide sophisticated services that take customer emotions into consideration while strictly protecting personal information, thereby maintaining a competitive advantage in the market.
[0389] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0390] Step 1:
[0391] The user records audio using a smartphone or dedicated device. The input is the user's voice, and the output is an audio data file. The microphone on the device captures the sound and saves it in a file format.
[0392] Step 2:
[0393] The terminal sends recorded audio data to the server. The input is an audio data file, and the output is data transfer to the server. The terminal encrypts the data using a secure protocol and sends it to the specified server address.
[0394] Step 3:
[0395] The server saves the received audio data to a database. The input is the audio data transferred from the terminal, and the output is a file stored in the database. The server verifies the integrity of the data before storing it in the database.
[0396] Step 4:
[0397] The server converts audio data into text information using a speech conversion unit. The input is an audio data file, and the output is text data. Speech recognition software analyzes the audio into text and generates text information.
[0398] Step 5:
[0399] The server uses a personal information identification unit to identify personal information from text data. The input is text data, and the output is the detected personal information. Natural language processing technology is used to extract names, addresses, and other information.
[0400] Step 6:
[0401] The server uses the information protection unit to encrypt the identified personal information. The input is the detected personal information, and the output is the encrypted personal information. An encryption algorithm is applied to ensure security.
[0402] Step 7:
[0403] The server uses an emotion analysis unit to estimate emotional states from voice data. The input is voice data, and the output is a classified emotional state. It analyzes the tone and speed of the voice to estimate emotions such as "joy" or "sadness."
[0404] Step 8:
[0405] The server transmits encrypted information and emotional states to an external service via a data communication unit. The input is encrypted personal information and emotional states, and the output is data transfer to an external service. Data is securely shared using an API.
[0406] (Application Example 2)
[0407] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0408] In today's world, there is a need to accurately monitor emotional changes while ensuring individual safety. In particular, there is a need for a system that can proactively detect and warn of risks associated with rapid emotional shifts. However, existing voice analysis technologies struggle to evaluate a user's emotional state in real time and generate appropriate alerts while maintaining the protection of personal information. Technologies are needed to overcome these limitations.
[0409] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0410] In this invention, the server includes means for acquiring acoustic information, means for converting acoustic information into text information, and means for analyzing the acoustic information to evaluate the user's emotional state. This makes it possible to detect emotional states from acoustic information and generate warnings for abnormal emotional changes while protecting personal information.
[0411] The "data acquisition unit" is a component that has the function of collecting acoustic information from the user's environment.
[0412] The "speech recognition unit" is a component that appropriately converts acquired acoustic information into text information and makes it into an analyzable format.
[0413] The "personal information detection unit" is the part of the system that has the function of identifying information that can identify an individual from textual information and analyzing it securely.
[0414] The "encryption processing unit" is the part of the system that has the function of encrypting identified personal information to protect it from unauthorized access by others.
[0415] The "emotion analysis unit" is a dedicated component that analyzes the tone and rhythm of the user's voice from acoustic information and evaluates their emotional state.
[0416] The "data transmission unit" is a configuration that has the function of transmitting processed data and encrypted information to external functions or services.
[0417] The "warning generation unit" is a component that has the function of issuing an appropriate warning when the user's emotional state exceeds a predetermined threshold.
[0418] To implement this invention, a system is needed that has a series of processes for collecting acoustic information, analyzing that information, and evaluating the user's emotional state. First, acoustic information is collected from the user's surroundings using a terminal such as a smartphone or a dedicated device. This data is securely transmitted to a server through a data acquisition unit. On the server, speech recognition software (e.g., Google Speech-to-Text API) is used to convert the acoustic information into text information.
[0419] The server analyzes textual information using natural language processing technology, and the personal information identified by the personal information detection unit is protected by the encryption processing unit. Meanwhile, the emotion analysis unit uses a generative AI model (e.g., a custom model using TensorFlow) to evaluate the user's emotional state from acoustic information. This process includes the analysis of sound tone and rhythm, and emotions are classified into multiple categories.
[0420] If the emotional state exceeds a set threshold, the server sends an alert to emergency contacts via the warning generation unit. This procedure enables real-time emotion monitoring and warning generation based on acoustic information while maintaining the security of personal information.
[0421] For example, if a user is experiencing significant stress at home, the system immediately detects this change in emotion and sends a warning message stating, "The user may be experiencing stress." This warning is then sent via a smartphone app to pre-configured emergency contacts.
[0422] An example of a prompt message would be: "Use this prompt message to monitor the user's emotional changes. Analyze the emotional information obtained from the voice data and issue appropriate actions according to the urgency." This allows the system to effectively track the user's emotional state and take necessary actions.
[0423] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0424] Step 1:
[0425] The terminal collects acoustic information from the user's surroundings using its built-in microphone. This acoustic information becomes the input data. The terminal then prepares to transmit this data to the server via the data acquisition unit.
[0426] Step 2:
[0427] The server receives the acoustic information transmitted from the terminal in its data acquisition unit. Next, it uses speech recognition software (e.g., Google Speech-to-Text API) to convert this acoustic information into text information. This conversion results in the text information becoming the output data.
[0428] Step 3:
[0429] The server analyzes the text information converted by the speech recognition unit using natural language processing technology, and extracts personally identifiable information in the personal information detection unit. This personal information is the output of the processing and is then securely protected by the encryption processing unit.
[0430] Step 4:
[0431] The server uses an emotion analysis unit to analyze the tone and rhythm of the voice from the acoustic information, and utilizes a generative AI model to evaluate the user's emotional state. In this process, the acoustic information is the input data, and the emotion evaluation result is obtained as output data.
[0432] Step 5:
[0433] The server determines whether the emotional state has exceeded a set threshold based on the emotion evaluation results obtained from the emotion analysis unit. If the threshold is exceeded, the warning generation unit sends an alert to the emergency contact. At this time, the warning alert is output using the emotion evaluation results as input.
[0434] Step 6:
[0435] The user receives an alert from the server and takes the appropriate action according to the instructions in the prompt. In this step, the action based on the prompt appears as output.
[0436] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0437] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0438] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0439] [Third Embodiment]
[0440] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0441] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0442] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0443] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0444] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0445] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0446] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0447] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0448] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0449] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0450] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0451] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0452] The system of the present invention securely analyzes received audio data, minimizing the risk of personal information leakage while enabling integration with cloud services.
[0453] The system begins with the user recording calls and voice information and saving it as audio data on their device. This audio data is then sent to the server using a secure communication protocol.
[0454] The server stores the received audio data and converts it into text data using a speech recognition unit. This conversion makes the audio information searchable in text format, facilitating subsequent processing. Next, the server uses a personal information detection unit to identify personal information from the converted text data. Here, natural language processing technology is used to effectively extract important information such as names, addresses, phone numbers, and email addresses.
[0455] The identified personal information is encrypted by the server's encryption processing unit. This encryption protects the data when it is transmitted externally, preventing information leaks due to unauthorized access. The encrypted data is sent from the server to an external cloud service, where various analyses and processing are performed.
[0456] The results obtained from this external processing are returned to the server, which decrypts the data as needed and restores personal information. At this time, the decrypted information is kept only in a secure environment within the company, thus ensuring the protection of privacy.
[0457] As a concrete example, consider a case where a user records customer phone calls and uses a cloud service to analyze them. The server accurately transcribes the calls into text, masks customer personal information, and enables analysis of how customer interactions were conducted on the cloud. Based on the analysis results, specific improvement measures can then be fed back to the company.
[0458] This model allows companies to maximize the value of their data while complying with the law, enabling them to achieve both efficient operations and improved customer satisfaction.
[0459] The following describes the processing flow.
[0460] Step 1:
[0461] Users record phone calls or voice recordings and save the audio data to their devices. This audio data includes interactions and conversations with customers.
[0462] Step 2:
[0463] The terminal transmits voice data to the server using a secure communication protocol. Encryption is used during transmission to maintain data consistency and confidentiality.
[0464] Step 3:
[0465] The server stores the received audio data in its internal storage. When storing the data, it also stores metadata such as data identifiers and timestamps.
[0466] Step 4:
[0467] The server's speech recognition unit converts the received audio data into text data. In this process, a speech recognition algorithm is used to extract language from the audio file and convert it into the corresponding text.
[0468] Step 5:
[0469] The server's personal information detection unit analyzes the converted text data and uses internal natural language processing technology to identify personal information such as names, phone numbers, and addresses. The identified personal information is tagged and marked up for encryption.
[0470] Step 6:
[0471] The server's encryption processing unit encrypts the personal information marked up by the personal information detection unit. By using an encryption algorithm, the confidential information is converted into a format that cannot be identified.
[0472] Step 7:
[0473] The server's data transmission unit sends encrypted text data to an external cloud service. Since the data already has personal information protected, it is in a format suitable for external processing.
[0474] Step 8:
[0475] Cloud services receive encrypted data sent from servers and perform analysis and processing. For example, this can be used for evaluating customer service quality and analyzing trends.
[0476] Step 9:
[0477] The server receives processing results from the cloud service. The received result data is then fully utilized internally as needed.
[0478] Step 10:
[0479] The server's decryption unit decrypts encrypted data internally as needed. This process allows for the complete reconstruction of personal information only within the system, enabling detailed data analysis while protecting privacy.
[0480] (Example 1)
[0481] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0482] There is a need for a method to securely digitize audio information while preventing the leakage of personal information, and to efficiently analyze and utilize the data results on external platforms. However, existing technologies make it difficult to balance information security and effective data utilization, and the risk of personal information leakage is a particular concern.
[0483] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0484] In this invention, the server includes means for storing voice information with a data acquisition unit, means for converting the voice information into text information with a voice conversion unit, and means for converting the extracted important information into data with an information protection unit. This makes it possible to utilize advanced analysis on external platforms while ensuring the protection of personal information.
[0485] The term "data acquisition unit" refers to a device or function for collecting and storing audio information.
[0486] The term "voice conversion unit" refers to a device or function that converts voice information into digital data and then into text information.
[0487] The term "information detection unit" refers to a device or function used to identify and extract important data from textual information.
[0488] The "information protection unit" refers to the devices and functions that transform extracted sensitive data in order to securely protect it.
[0489] The "data transmission unit" refers to the device or function used to transmit the converted data to an external platform.
[0490] An "external platform" refers to an external computing environment that can analyze data received by the server and return the results.
[0491] The term "restoration unit" refers to a device or function that restores result data received from an external platform to its original state as needed.
[0492] "Generative algorithms" refer to methods used in data analysis to generate new insights and solutions.
[0493] This invention provides a system that processes voice information securely and efficiently, enabling analysis on external platforms while protecting personal information. The following describes embodiments for implementing this system.
[0494] First, the user records the call or audio information. The device saves this audio information in an appropriate digital format (e.g., WAV or MP3). When the device sends the saved audio information to the server, it uses a secure protocol such as TLS or SSL.
[0495] When the server receives audio information, it converts it into text information using a speech conversion unit. This conversion utilizes a speech recognition service (e.g., a general-purpose speech recognition API). The converted text information is then analyzed by an information detection unit to extract important data (e.g., name and address). Accuracy can be improved by utilizing natural language processing techniques at this stage.
[0496] The identified sensitive information is transformed by the Information Protection Unit into a secure, transmittable format. This information is then transmitted to an external platform by the Data Transmission Unit, where it is analyzed using generative algorithms. Through this analysis, users can gain new insights and improvement strategies.
[0497] As a concrete example, consider a case where a user records phone calls with customers and uses an external service to derive measures to improve customer satisfaction. The server transcribes the calls into text, analyzes them while protecting personal information, and then provides improvement suggestions based on this analysis.
[0498] An example of a prompt might be the instruction, "Analyze customer service calls and generate improvement suggestions." Based on this prompt, a generative AI model generates useful suggestions from the data. In this way, it becomes possible to maximize the use of corporate data while ensuring security, and to support efficient and effective decision-making.
[0499] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0500] Step 1:
[0501] The user records phone calls and audio data. The input is the audio from the call. The device saves this audio in a digital format (e.g., a WAV file). The output is the saved digital audio file.
[0502] Step 2:
[0503] The terminal sends the stored audio file to the server. The input is a digital audio file. The terminal uses secure protocols such as TLS or SSL to ensure secure data transmission. The output is the audio file that has been securely delivered to the server.
[0504] Step 3:
[0505] The server processes the received audio file in its speech conversion unit and converts it into text information. The input is an audio data file. The server utilizes a speech recognition API to convert digital audio into text. The output is the converted text data.
[0506] Step 4:
[0507] The server analyzes textual information, and the information detection unit identifies personal information. The input is converted text data. Important information (e.g., name, address) is extracted using natural language processing technology. The output is the identified personal information.
[0508] Step 5:
[0509] The server encrypts and transforms the identified personal information using the information protection unit. The input is identified personal information. The server transforms the personal information using an encryption algorithm (e.g., AES). The output is the encrypted personal information.
[0510] Step 6:
[0511] The server transmits encrypted personal information to an external platform. The input is encrypted personal information. This data is securely transmitted to the external platform via the data transmission unit. The output is the data transmitted to the external platform.
[0512] Step 7:
[0513] The external platform analyzes data using a generative AI model and generates results. The input is encrypted personal information. The external system performs the analysis and generates new insights and suggestions. The output is the analysis results and suggested actions.
[0514] Step 8:
[0515] The server receives analysis results from an external platform and performs decryption. The input is the analysis results sent from the external platform. The server uses a recovery unit to decrypt the results as needed and return them to a usable format. The output is the decrypted analysis results.
[0516] Step 9:
[0517] Users or internal personnel take action to make decisions and implement improvements based on the decoded results. The input is the decoded analysis findings and suggestions. The output is the actions taken and the results obtained therefrom.
[0518] (Application Example 1)
[0519] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0520] In today's online commerce environment, streamlining customer service is a crucial challenge. However, analyzing customer call content carries the risk of personal information leaks, requiring appropriate protective measures. Furthermore, rapidly and accurately analyzing large amounts of call data and deriving concrete improvement measures is a technically challenging task.
[0521] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0522] In this invention, the server includes a data acquisition unit for acquiring voice information, a voice conversion unit for converting the voice information into text data, and an information detection unit for extracting personal information from the text data. This makes it possible to efficiently analyze customer call data while protecting personal information and use the results to improve customer service.
[0523] The "data acquisition unit" is a device that collects voice information and incorporates it as initial data.
[0524] The "voice conversion unit" is a device that converts collected voice information into text data.
[0525] The "information detection unit" is a device that extracts personal information from text data and identifies the necessary information.
[0526] The "processing unit" is a device that encrypts the extracted personal information to ensure data security.
[0527] The "transmission unit" is a device used to transmit encrypted information to an external system.
[0528] The "analysis unit" is a device that uses analysis results obtained from external systems to improve internal operational efficiency.
[0529] The "decryption unit" is a device that restores encrypted personal information and reconstructs the original information in a secure environment.
[0530] This invention realizes a system that efficiently acquires and analyzes voice information and utilizes the results. First, the user makes a call with a customer through smart glasses. Voice information is collected by the data acquisition unit. The smart glasses can be wearable devices available on the market, such as Google Glass.
[0531] The collected audio information is converted into text data through a speech-to-text unit. This process can utilize speech recognition technologies such as the Google Speech-to-Text API. The converted text data is then processed by an information detection unit, which uses natural language processing techniques to identify personal information. NLP tools such as SpaCy are useful in this step.
[0532] Personal information identified by the information detection unit is encrypted by the processing unit to ensure security. Encryption libraries such as OpenSSL can be used for this encryption. The encrypted information is then transmitted to an analysis service on the cloud via the transmission unit.
[0533] On the cloud side, data is analyzed, and the results are fed back internally through the analysis unit. For example, areas for improvement in customer service are identified. The analysis unit can use a generative AI model to generate prompt messages and suggest improvements.
[0534] As a concrete example, consider a case where a customer inquires about "order cancellation" at a support center. This inquiry is recorded as audio, converted to text, and then keywords such as "order" and "cancellation" are extracted by the information detection unit. As a result, if similar inquiries occur frequently, common patterns in the reasons for cancellation can be identified, which can be used to improve support operations.
[0535] An example of a prompt for the generating AI model could be text such as, "Analyze the patterns in why customers cancel orders." This allows the system to efficiently analyze data and generate specific improvement suggestions.
[0536] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0537] Step 1:
[0538] The user makes phone calls with customers using smart glasses. The calls are collected as audio information by the data acquisition unit. The audio information is stored as raw data in the smart glasses' memory.
[0539] Step 2:
[0540] The terminal sends the collected audio information to the server. The server uses a speech conversion unit to convert the audio information into text data. This process utilizes speech recognition software (e.g., Google Speech-to-Text API). The audio data (input) is converted into text data (output).
[0541] Step 3:
[0542] The server processes the text data in its information detection unit. At this stage, natural language processing technology (e.g., SpaCy) is used to identify personal information. Personal information such as names and addresses (output) is extracted from the text data (input).
[0543] Step 4:
[0544] The server encrypts the identified personal information in its processing unit. Using an encryption library (e.g., OpenSSL), the personal information (input) is converted into secure encrypted data (output).
[0545] Step 5:
[0546] The server sends encrypted data to the cloud analysis service via the transmission unit. The cloud system performs the analysis and generates prompt messages using an AI model that generates the results. The encrypted data (input) is converted into analysis results (output), and prompt messages are generated.
[0547] Step 6:
[0548] The analysis results returned from the cloud are fed back by the analysis unit on the server. The server then identifies specific areas for improvement to enhance internal operational efficiency. The analysis results (input) are then used as actual improvement suggestions (output).
[0549] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0550] The system in this invention has advanced data processing capabilities that simultaneously evaluate the user's emotional state in addition to analyzing voice data. This invention aims to improve communication with customers by ensuring the secure handling of personal information and enabling data analysis that takes user emotions into consideration.
[0551] The system operation begins with the user recording audio data. The user uses a smartphone or dedicated device to record phone calls and conversations, saving the data to the device. The recorded audio data is then transmitted from the device to the server using a secure protocol.
[0552] The server stores the received audio data in a database and simultaneously converts it into text data using a speech recognition unit. This conversion makes the audio information machine-processable. When analyzing the text data, the server utilizes a personal information detection unit to identify personal information such as names and addresses, and securely protects this information using an encryption processing unit.
[0553] Furthermore, this system is equipped with an emotion engine that estimates the user's emotions from voice data. The emotion engine analyzes the tone, speed, and rhythm of the voice and classifies the user's emotional state into several categories (e.g., joy, sadness, anger). This information is stored in a database as an essential element for improving customer service.
[0554] Encrypted data and analysis results from the emotion engine are securely transmitted to an external cloud service. Personal information is masked during this process, enabling in-depth analysis by the external service. Once the external processing results are returned to the server, the system decrypts the data as needed and uses that information to develop detailed customer service improvement strategies internally.
[0555] For example, when a recording of a customer's phone call to customer support is analyzed, the customer's emotions can be identified as "dissatisfaction" from the recording, and the analysis results are provided while appropriately protecting the personal information of the call. This allows the support team to take concrete approaches to improve the quality of customer service.
[0556] This model allows companies to protect personal information while providing sophisticated services that take customer emotions into consideration, thereby securing a competitive advantage in the market.
[0557] The following describes the processing flow.
[0558] Step 1:
[0559] The user records audio data. Using a smartphone or dedicated device, calls and conversations are recorded in real time.
[0560] Step 2:
[0561] The device transmits the recorded audio data to the server via a secure protocol. The data is encrypted during transmission to maintain confidentiality.
[0562] Step 3:
[0563] The server stores the received audio data in a database. The data is also accompanied by metadata such as an identifier and the date and time of recording.
[0564] Step 4:
[0565] The server's speech recognition unit converts the audio data into text data. The converted text becomes the basic data for analyzing the conversation content as written text.
[0566] Step 5:
[0567] The server's personal information detection unit analyzes text data and uses natural language processing technology to identify personal information such as names and addresses. The identified information is then marked for encryption.
[0568] Step 6:
[0569] The server's encryption processing unit encrypts the identified personal information. The encrypted personal information is rendered harmless, ensuring security when exchanging data with external parties.
[0570] Step 7:
[0571] The server's emotion engine analyzes the voice data and estimates the user's emotional state from the tone and rhythm of their voice. The estimated emotion is then classified into categories such as "joy" or "dissatisfaction."
[0572] Step 8:
[0573] The server sends encrypted text data and sentiment analysis results to an external cloud service. This external service analyzes the data and generates detailed insights.
[0574] Step 9:
[0575] The server receives analysis results returned from an external service. These results may include, for example, customer sentiment trends and areas for improvement in service delivery.
[0576] Step 10:
[0577] The server's decryption unit decrypts the data as needed. This is used for internal verification and specific customer support. This ensures that the information can be used while maintaining security.
[0578] (Example 2)
[0579] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0580] In analyzing voice data, it is necessary to accurately estimate the user's emotional state, securely protect personal information, and quickly and efficiently integrate with external services. Conventional systems have challenges in the accuracy of emotion estimation and the protection of personal information, and rapid external integration is required.
[0581] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0582] In this invention, the server includes a data acquisition unit for acquiring voice information, a voice conversion unit for converting the voice information into text information, a personal information identification unit for identifying personal information from the text information, an emotion analysis unit for estimating emotional states from voice information, an information protection unit for encrypting identified personal information, and a data communication unit for transmitting the encrypted information and emotional states to an external service. This enables data analysis that takes into account the user's emotional state, secure protection of personal information, and rapid external collaboration.
[0583] A "data acquisition unit" is a device or function that collects voice information and incorporates it into the system.
[0584] A "speech conversion unit" refers to a device or technology that has the function of analyzing acquired speech information and converting it into text information.
[0585] A "personal information identification unit" is a device or process that has the function of accurately extracting and identifying information related to a specific individual from textual information.
[0586] "Information protection" refers to a device or technology that has the function of encrypting identified personal information in order to securely protect that information.
[0587] A "sentiment analysis unit" refers to a device or software that has the function of estimating and classifying the emotional state of a speaker by analyzing audio information.
[0588] A "data communication unit" refers to a device or function equipped with communication means for transmitting processed information to an external service.
[0589] The system based on this invention analyzes voice information through collaboration between the user, terminal, and server, securely protects personal information, evaluates the user's emotional state, and integrates with external services. The system aims to improve service quality and facilitate smooth communication with customers by ensuring that each department operates appropriately.
[0590] First, the user records audio via their smartphone or a dedicated device. The recorded audio data is then sent to the server by the device using a secure protocol. During this process, the device encrypts the audio data to ensure its security.
[0591] When the server receives data, it first stores it in a database and then uses a speech conversion unit (e.g., speech recognition software) to convert the speech data into text information. Next, it uses a personal information identification unit to identify personal information from the text information. At this stage, natural language processing technology is utilized, and the identified personal information is encrypted by the information protection unit.
[0592] Next, the emotion analysis unit installed on the server analyzes the voice data and infers the emotional state. The emotion engine classifies the information into categories such as "joy" and "sadness" based on the tone and speed of the voice. The results are stored in a database and used for marketing and customer service.
[0593] Furthermore, the server securely transmits encrypted personal information and sentiment analysis data to an external cloud service via the data communication unit. This enables detailed analysis by the external service, and the analysis results are sent back to the server for internal use.
[0594] As a concrete example, consider a scenario where a customer calls customer support and the conversation is recorded. In this situation, the system recognizes the emotion of the call as "dissatisfaction" and provides analysis results while protecting the personal information of the caller. In this way, the support team can develop specific measures to improve the quality of customer service.
[0595] Examples of prompt statements in generative AI models include the following:
[0596] "We estimate customer emotions based on voice data and propose specific countermeasures."
[0597] Such systems enable companies to provide sophisticated services that take customer emotions into consideration while strictly protecting personal information, thereby maintaining a competitive advantage in the market.
[0598] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0599] Step 1:
[0600] The user records audio using a smartphone or dedicated device. The input is the user's voice, and the output is an audio data file. The microphone on the device captures the sound and saves it in a file format.
[0601] Step 2:
[0602] The terminal sends recorded audio data to the server. The input is an audio data file, and the output is data transfer to the server. The terminal encrypts the data using a secure protocol and sends it to the specified server address.
[0603] Step 3:
[0604] The server saves the received audio data to a database. The input is the audio data transferred from the terminal, and the output is a file stored in the database. The server verifies the integrity of the data before storing it in the database.
[0605] Step 4:
[0606] The server converts audio data into text information using a speech conversion unit. The input is an audio data file, and the output is text data. Speech recognition software analyzes the audio into text and generates text information.
[0607] Step 5:
[0608] The server uses a personal information identification unit to identify personal information from text data. The input is text data, and the output is the detected personal information. Natural language processing technology is used to extract names, addresses, and other information.
[0609] Step 6:
[0610] The server uses the information protection unit to encrypt the identified personal information. The input is the detected personal information, and the output is the encrypted personal information. An encryption algorithm is applied to ensure security.
[0611] Step 7:
[0612] The server uses an emotion analysis unit to estimate emotional states from voice data. The input is voice data, and the output is a classified emotional state. It analyzes the tone and speed of the voice to estimate emotions such as "joy" or "sadness."
[0613] Step 8:
[0614] The server transmits encrypted information and emotional states to an external service via a data communication unit. The input is encrypted personal information and emotional states, and the output is data transfer to an external service. Data is securely shared using an API.
[0615] (Application Example 2)
[0616] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0617] In today's world, there is a need to accurately monitor emotional changes while ensuring individual safety. In particular, there is a need for a system that can proactively detect and warn of risks associated with rapid emotional shifts. However, existing voice analysis technologies struggle to evaluate a user's emotional state in real time and generate appropriate alerts while maintaining the protection of personal information. Technologies are needed to overcome these limitations.
[0618] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0619] In this invention, the server includes means for acquiring acoustic information, means for converting acoustic information into text information, and means for analyzing the acoustic information to evaluate the user's emotional state. This makes it possible to detect emotional states from acoustic information and generate warnings for abnormal emotional changes while protecting personal information.
[0620] The "data acquisition unit" is a component that has the function of collecting acoustic information from the user's environment.
[0621] The "speech recognition unit" is a component that appropriately converts acquired acoustic information into text information and makes it into an analyzable format.
[0622] The "personal information detection unit" is the part of the system that has the function of identifying information that can identify an individual from textual information and analyzing it securely.
[0623] The "encryption processing unit" is the part of the system that has the function of encrypting identified personal information to protect it from unauthorized access by others.
[0624] The "emotion analysis unit" is a dedicated component that analyzes the tone and rhythm of the user's voice from acoustic information and evaluates their emotional state.
[0625] The "data transmission unit" is a configuration that has the function of transmitting processed data and encrypted information to external functions or services.
[0626] The "warning generation unit" is a component that has the function of issuing an appropriate warning when the user's emotional state exceeds a predetermined threshold.
[0627] To implement this invention, a system is needed that has a series of processes for collecting acoustic information, analyzing that information, and evaluating the user's emotional state. First, acoustic information is collected from the user's surroundings using a terminal such as a smartphone or a dedicated device. This data is securely transmitted to a server through a data acquisition unit. On the server, speech recognition software (e.g., Google Speech-to-Text API) is used to convert the acoustic information into text information.
[0628] The server analyzes textual information using natural language processing technology, and the personal information identified by the personal information detection unit is protected by the encryption processing unit. Meanwhile, the emotion analysis unit uses a generative AI model (e.g., a custom model using TensorFlow) to evaluate the user's emotional state from acoustic information. This process includes the analysis of sound tone and rhythm, and emotions are classified into multiple categories.
[0629] If the emotional state exceeds a set threshold, the server sends an alert to emergency contacts via the warning generation unit. This procedure enables real-time emotion monitoring and warning generation based on acoustic information while maintaining the security of personal information.
[0630] For example, if a user is experiencing significant stress at home, the system immediately detects this change in emotion and sends a warning message stating, "The user may be experiencing stress." This warning is then sent via a smartphone app to pre-configured emergency contacts.
[0631] An example of a prompt message would be: "Use this prompt message to monitor the user's emotional changes. Analyze the emotional information obtained from the voice data and issue appropriate actions according to the urgency." This allows the system to effectively track the user's emotional state and take necessary actions.
[0632] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0633] Step 1:
[0634] The terminal collects acoustic information from the user's surroundings using its built-in microphone. This acoustic information becomes the input data. The terminal then prepares to transmit this data to the server via the data acquisition unit.
[0635] Step 2:
[0636] The server receives the acoustic information transmitted from the terminal in its data acquisition unit. Next, it uses speech recognition software (e.g., Google Speech-to-Text API) to convert this acoustic information into text information. This conversion results in the text information becoming the output data.
[0637] Step 3:
[0638] The server analyzes the text information converted by the speech recognition unit using natural language processing technology, and extracts personally identifiable information in the personal information detection unit. This personal information is the output of the processing and is then securely protected by the encryption processing unit.
[0639] Step 4:
[0640] The server uses an emotion analysis unit to analyze the tone and rhythm of the voice from the acoustic information, and utilizes a generative AI model to evaluate the user's emotional state. In this process, the acoustic information is the input data, and the emotion evaluation result is obtained as output data.
[0641] Step 5:
[0642] The server determines whether the emotional state has exceeded a set threshold based on the emotion evaluation results obtained from the emotion analysis unit. If the threshold is exceeded, the warning generation unit sends an alert to the emergency contact. At this time, the warning alert is output using the emotion evaluation results as input.
[0643] Step 6:
[0644] The user receives an alert from the server and takes the appropriate action according to the instructions in the prompt. In this step, the action based on the prompt appears as output.
[0645] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0646] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0647] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0648] [Fourth Embodiment]
[0649] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0650] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0651] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0652] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0653] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0654] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0655] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0656] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0657] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0658] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0659] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0660] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0661] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0662] The system of the present invention securely analyzes received audio data, minimizing the risk of personal information leakage while enabling integration with cloud services.
[0663] The system begins with the user recording calls and voice information and saving it as audio data on their device. This audio data is then sent to the server using a secure communication protocol.
[0664] The server stores the received audio data and converts it into text data using a speech recognition unit. This conversion makes the audio information searchable in text format, facilitating subsequent processing. Next, the server uses a personal information detection unit to identify personal information from the converted text data. Here, natural language processing technology is used to effectively extract important information such as names, addresses, phone numbers, and email addresses.
[0665] The identified personal information is encrypted by the server's encryption processing unit. This encryption protects the data when it is transmitted externally, preventing information leaks due to unauthorized access. The encrypted data is sent from the server to an external cloud service, where various analyses and processing are performed.
[0666] The results obtained from this external processing are returned to the server, which decrypts the data as needed and restores personal information. At this time, the decrypted information is kept only in a secure environment within the company, thus ensuring the protection of privacy.
[0667] As a concrete example, consider a case where a user records customer phone calls and uses a cloud service to analyze them. The server accurately transcribes the calls into text, masks customer personal information, and enables analysis of how customer interactions were conducted on the cloud. Based on the analysis results, specific improvement measures can then be fed back to the company.
[0668] This model allows companies to maximize the value of their data while complying with the law, enabling them to achieve both efficient operations and improved customer satisfaction.
[0669] The following describes the processing flow.
[0670] Step 1:
[0671] Users record phone calls or voice recordings and save the audio data to their devices. This audio data includes interactions and conversations with customers.
[0672] Step 2:
[0673] The terminal transmits voice data to the server using a secure communication protocol. Encryption is used during transmission to maintain data consistency and confidentiality.
[0674] Step 3:
[0675] The server stores the received audio data in its internal storage. When storing the data, it also stores metadata such as data identifiers and timestamps.
[0676] Step 4:
[0677] The server's speech recognition unit converts the received audio data into text data. In this process, a speech recognition algorithm is used to extract language from the audio file and convert it into the corresponding text.
[0678] Step 5:
[0679] The server's personal information detection unit analyzes the converted text data and uses internal natural language processing technology to identify personal information such as names, phone numbers, and addresses. The identified personal information is tagged and marked up for encryption.
[0680] Step 6:
[0681] The server's encryption processing unit encrypts the personal information marked up by the personal information detection unit. By using an encryption algorithm, the confidential information is converted into a format that cannot be identified.
[0682] Step 7:
[0683] The server's data transmission unit sends encrypted text data to an external cloud service. Since the data already has personal information protected, it is in a format suitable for external processing.
[0684] Step 8:
[0685] Cloud services receive encrypted data sent from servers and perform analysis and processing. For example, this can be used for evaluating customer service quality and analyzing trends.
[0686] Step 9:
[0687] The server receives processing results from the cloud service. The received result data is then fully utilized internally as needed.
[0688] Step 10:
[0689] The server's decryption unit decrypts encrypted data internally as needed. This process allows for the complete reconstruction of personal information only within the system, enabling detailed data analysis while protecting privacy.
[0690] (Example 1)
[0691] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0692] There is a need for a method to securely digitize audio information while preventing the leakage of personal information, and to efficiently analyze and utilize the data results on external platforms. However, existing technologies make it difficult to balance information security and effective data utilization, and the risk of personal information leakage is a particular concern.
[0693] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0694] In this invention, the server includes means for storing voice information with a data acquisition unit, means for converting the voice information into text information with a voice conversion unit, and means for converting the extracted important information into data with an information protection unit. This makes it possible to utilize advanced analysis on external platforms while ensuring the protection of personal information.
[0695] The term "data acquisition unit" refers to a device or function for collecting and storing audio information.
[0696] The term "voice conversion unit" refers to a device or function that converts voice information into digital data and then into text information.
[0697] The term "information detection unit" refers to a device or function used to identify and extract important data from textual information.
[0698] The "information protection unit" refers to the devices and functions that transform extracted sensitive data in order to securely protect it.
[0699] The "data transmission unit" refers to the device or function used to transmit the converted data to an external platform.
[0700] An "external platform" refers to an external computing environment that can analyze data received by the server and return the results.
[0701] The term "restoration unit" refers to a device or function that restores result data received from an external platform to its original state as needed.
[0702] "Generative algorithms" refer to methods used in data analysis to generate new insights and solutions.
[0703] This invention provides a system that processes voice information securely and efficiently, enabling analysis on external platforms while protecting personal information. The following describes embodiments for implementing this system.
[0704] First, the user records the call or audio information. The device saves this audio information in an appropriate digital format (e.g., WAV or MP3). When the device sends the saved audio information to the server, it uses a secure protocol such as TLS or SSL.
[0705] When the server receives audio information, it converts it into text information using a speech conversion unit. This conversion utilizes a speech recognition service (e.g., a general-purpose speech recognition API). The converted text information is then analyzed by an information detection unit to extract important data (e.g., name and address). Accuracy can be improved by utilizing natural language processing techniques at this stage.
[0706] The identified sensitive information is transformed by the Information Protection Unit into a secure, transmittable format. This information is then transmitted to an external platform by the Data Transmission Unit, where it is analyzed using generative algorithms. Through this analysis, users can gain new insights and improvement strategies.
[0707] As a concrete example, consider a case where a user records phone calls with customers and uses an external service to derive measures to improve customer satisfaction. The server transcribes the calls into text, analyzes them while protecting personal information, and then provides improvement suggestions based on this analysis.
[0708] An example of a prompt might be the instruction, "Analyze customer service calls and generate improvement suggestions." Based on this prompt, a generative AI model generates useful suggestions from the data. In this way, it becomes possible to maximize the use of corporate data while ensuring security, and to support efficient and effective decision-making.
[0709] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0710] Step 1:
[0711] The user records phone calls and audio data. The input is the audio from the call. The device saves this audio in a digital format (e.g., a WAV file). The output is the saved digital audio file.
[0712] Step 2:
[0713] The terminal sends the stored audio file to the server. The input is a digital audio file. The terminal uses secure protocols such as TLS or SSL to ensure secure data transmission. The output is the audio file that has been securely delivered to the server.
[0714] Step 3:
[0715] The server processes the received audio file in its speech conversion unit and converts it into text information. The input is an audio data file. The server utilizes a speech recognition API to convert digital audio into text. The output is the converted text data.
[0716] Step 4:
[0717] The server analyzes textual information, and the information detection unit identifies personal information. The input is converted text data. Important information (e.g., name, address) is extracted using natural language processing technology. The output is the identified personal information.
[0718] Step 5:
[0719] The server encrypts and transforms the identified personal information using the information protection unit. The input is identified personal information. The server transforms the personal information using an encryption algorithm (e.g., AES). The output is the encrypted personal information.
[0720] Step 6:
[0721] The server transmits encrypted personal information to an external platform. The input is encrypted personal information. This data is securely transmitted to the external platform via the data transmission unit. The output is the data transmitted to the external platform.
[0722] Step 7:
[0723] The external platform analyzes data using a generative AI model and generates results. The input is encrypted personal information. The external system performs the analysis and generates new insights and suggestions. The output is the analysis results and suggested actions.
[0724] Step 8:
[0725] The server receives analysis results from an external platform and performs decryption. The input is the analysis results sent from the external platform. The server uses a recovery unit to decrypt the results as needed and return them to a usable format. The output is the decrypted analysis results.
[0726] Step 9:
[0727] Users or internal personnel take action to make decisions and implement improvements based on the decoded results. The input is the decoded analysis findings and suggestions. The output is the actions taken and the results obtained therefrom.
[0728] (Application Example 1)
[0729] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0730] In today's online commerce environment, streamlining customer service is a crucial challenge. However, analyzing customer call content carries the risk of personal information leaks, requiring appropriate protective measures. Furthermore, rapidly and accurately analyzing large amounts of call data and deriving concrete improvement measures is a technically challenging task.
[0731] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0732] In this invention, the server includes a data acquisition unit for acquiring voice information, a voice conversion unit for converting the voice information into text data, and an information detection unit for extracting personal information from the text data. This makes it possible to efficiently analyze customer call data while protecting personal information and use the results to improve customer service.
[0733] The "data acquisition unit" is a device that collects voice information and incorporates it as initial data.
[0734] The "voice conversion unit" is a device that converts collected voice information into text data.
[0735] The "information detection unit" is a device that extracts personal information from text data and identifies the necessary information.
[0736] The "processing unit" is a device that encrypts the extracted personal information to ensure data security.
[0737] The "transmission unit" is a device used to transmit encrypted information to an external system.
[0738] The "analysis unit" is a device that uses analysis results obtained from external systems to improve internal operational efficiency.
[0739] The "decryption unit" is a device that restores encrypted personal information and reconstructs the original information in a secure environment.
[0740] This invention realizes a system that efficiently acquires and analyzes voice information and utilizes the results. First, the user makes a call with a customer through smart glasses. Voice information is collected by the data acquisition unit. The smart glasses can be wearable devices available on the market, such as Google Glass.
[0741] The collected audio information is converted into text data through a speech-to-text unit. This process can utilize speech recognition technologies such as the Google Speech-to-Text API. The converted text data is then processed by an information detection unit, which uses natural language processing techniques to identify personal information. NLP tools such as SpaCy are useful in this step.
[0742] Personal information identified by the information detection unit is encrypted by the processing unit to ensure security. Encryption libraries such as OpenSSL can be used for this encryption. The encrypted information is then transmitted to an analysis service on the cloud via the transmission unit.
[0743] On the cloud side, data is analyzed, and the results are fed back internally through the analysis unit. For example, areas for improvement in customer service are identified. The analysis unit can use a generative AI model to generate prompt messages and suggest improvements.
[0744] As a concrete example, consider a case where a customer inquires about "order cancellation" at a support center. This inquiry is recorded as audio, converted to text, and then keywords such as "order" and "cancellation" are extracted by the information detection unit. As a result, if similar inquiries occur frequently, common patterns in the reasons for cancellation can be identified, which can be used to improve support operations.
[0745] An example of a prompt for the generating AI model could be text such as, "Analyze the patterns in why customers cancel orders." This allows the system to efficiently analyze data and generate specific improvement suggestions.
[0746] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0747] Step 1:
[0748] The user makes phone calls with customers using smart glasses. The calls are collected as audio information by the data acquisition unit. The audio information is stored as raw data in the smart glasses' memory.
[0749] Step 2:
[0750] The terminal sends the collected audio information to the server. The server uses a speech conversion unit to convert the audio information into text data. This process utilizes speech recognition software (e.g., Google Speech-to-Text API). The audio data (input) is converted into text data (output).
[0751] Step 3:
[0752] The server processes the text data in its information detection unit. At this stage, natural language processing technology (e.g., SpaCy) is used to identify personal information. Personal information such as names and addresses (output) is extracted from the text data (input).
[0753] Step 4:
[0754] The server encrypts the identified personal information in its processing unit. Using an encryption library (e.g., OpenSSL), the personal information (input) is converted into secure encrypted data (output).
[0755] Step 5:
[0756] The server sends encrypted data to the cloud analysis service via the transmission unit. The cloud system performs the analysis and generates prompt messages using an AI model that generates the results. The encrypted data (input) is converted into analysis results (output), and prompt messages are generated.
[0757] Step 6:
[0758] The analysis results returned from the cloud are fed back by the analysis unit on the server. The server then identifies specific areas for improvement to enhance internal operational efficiency. The analysis results (input) are then used as actual improvement suggestions (output).
[0759] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0760] The system in this invention has advanced data processing capabilities that simultaneously evaluate the user's emotional state in addition to analyzing voice data. This invention aims to improve communication with customers by ensuring the secure handling of personal information and enabling data analysis that takes user emotions into consideration.
[0761] The system operation begins with the user recording audio data. The user uses a smartphone or dedicated device to record phone calls and conversations, saving the data to the device. The recorded audio data is then transmitted from the device to the server using a secure protocol.
[0762] The server stores the received audio data in a database and simultaneously converts it into text data using a speech recognition unit. This conversion makes the audio information machine-processable. When analyzing the text data, the server utilizes a personal information detection unit to identify personal information such as names and addresses, and securely protects this information using an encryption processing unit.
[0763] Furthermore, this system is equipped with an emotion engine that estimates the user's emotions from voice data. The emotion engine analyzes the tone, speed, and rhythm of the voice and classifies the user's emotional state into several categories (e.g., joy, sadness, anger). This information is stored in a database as an essential element for improving customer service.
[0764] Encrypted data and analysis results from the emotion engine are securely transmitted to an external cloud service. Personal information is masked during this process, enabling in-depth analysis by the external service. Once the external processing results are returned to the server, the system decrypts the data as needed and uses that information to develop detailed customer service improvement strategies internally.
[0765] For example, when a recording of a customer's phone call to customer support is analyzed, the customer's emotions can be identified as "dissatisfaction" from the recording, and the analysis results are provided while appropriately protecting the personal information of the call. This allows the support team to take concrete approaches to improve the quality of customer service.
[0766] This model allows companies to protect personal information while providing sophisticated services that take customer emotions into consideration, thereby securing a competitive advantage in the market.
[0767] The following describes the processing flow.
[0768] Step 1:
[0769] The user records audio data. Using a smartphone or dedicated device, calls and conversations are recorded in real time.
[0770] Step 2:
[0771] The device transmits the recorded audio data to the server via a secure protocol. The data is encrypted during transmission to maintain confidentiality.
[0772] Step 3:
[0773] The server stores the received audio data in a database. The data is also accompanied by metadata such as an identifier and the date and time of recording.
[0774] Step 4:
[0775] The server's speech recognition unit converts the audio data into text data. The converted text becomes the basic data for analyzing the conversation content as written text.
[0776] Step 5:
[0777] The server's personal information detection unit analyzes text data and uses natural language processing technology to identify personal information such as names and addresses. The identified information is then marked for encryption.
[0778] Step 6:
[0779] The server's encryption processing unit encrypts the identified personal information. The encrypted personal information is rendered harmless, ensuring security when exchanging data with external parties.
[0780] Step 7:
[0781] The server's emotion engine analyzes the voice data and estimates the user's emotional state from the tone and rhythm of their voice. The estimated emotion is then classified into categories such as "joy" or "dissatisfaction."
[0782] Step 8:
[0783] The server sends encrypted text data and sentiment analysis results to an external cloud service. This external service analyzes the data and generates detailed insights.
[0784] Step 9:
[0785] The server receives analysis results returned from an external service. These results may include, for example, customer sentiment trends and areas for improvement in service delivery.
[0786] Step 10:
[0787] The server's decryption unit decrypts the data as needed. This is used for internal verification and specific customer support. This ensures that the information can be used while maintaining security.
[0788] (Example 2)
[0789] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0790] In analyzing voice data, it is necessary to accurately estimate the user's emotional state, securely protect personal information, and quickly and efficiently integrate with external services. Conventional systems have challenges in the accuracy of emotion estimation and the protection of personal information, and rapid external integration is required.
[0791] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0792] In this invention, the server includes a data acquisition unit for acquiring voice information, a voice conversion unit for converting the voice information into text information, a personal information identification unit for identifying personal information from the text information, an emotion analysis unit for estimating emotional states from voice information, an information protection unit for encrypting identified personal information, and a data communication unit for transmitting the encrypted information and emotional states to an external service. This enables data analysis that takes into account the user's emotional state, secure protection of personal information, and rapid external collaboration.
[0793] A "data acquisition unit" is a device or function that collects voice information and incorporates it into the system.
[0794] A "speech conversion unit" refers to a device or technology that has the function of analyzing acquired speech information and converting it into text information.
[0795] A "personal information identification unit" is a device or process that has the function of accurately extracting and identifying information related to a specific individual from textual information.
[0796] "Information protection" refers to a device or technology that has the function of encrypting identified personal information in order to securely protect that information.
[0797] A "sentiment analysis unit" refers to a device or software that has the function of estimating and classifying the emotional state of a speaker by analyzing audio information.
[0798] A "data communication unit" refers to a device or function equipped with communication means for transmitting processed information to an external service.
[0799] The system based on this invention analyzes voice information through collaboration between the user, terminal, and server, securely protects personal information, evaluates the user's emotional state, and integrates with external services. The system aims to improve service quality and facilitate smooth communication with customers by ensuring that each department operates appropriately.
[0800] First, the user records audio via their smartphone or a dedicated device. The recorded audio data is then sent to the server by the device using a secure protocol. During this process, the device encrypts the audio data to ensure its security.
[0801] When the server receives data, it first stores it in a database and then uses a speech conversion unit (e.g., speech recognition software) to convert the speech data into text information. Next, it uses a personal information identification unit to identify personal information from the text information. At this stage, natural language processing technology is utilized, and the identified personal information is encrypted by the information protection unit.
[0802] Next, the emotion analysis unit installed on the server analyzes the voice data and infers the emotional state. The emotion engine classifies the information into categories such as "joy" and "sadness" based on the tone and speed of the voice. The results are stored in a database and used for marketing and customer service.
[0803] Furthermore, the server securely transmits encrypted personal information and sentiment analysis data to an external cloud service via the data communication unit. This enables detailed analysis by the external service, and the analysis results are sent back to the server for internal use.
[0804] As a concrete example, consider a scenario where a customer calls customer support and the conversation is recorded. In this situation, the system recognizes the emotion of the call as "dissatisfaction" and provides analysis results while protecting the personal information of the caller. In this way, the support team can develop specific measures to improve the quality of customer service.
[0805] Examples of prompt statements in generative AI models include the following:
[0806] "We estimate customer emotions based on voice data and propose specific countermeasures."
[0807] Such systems enable companies to provide sophisticated services that take customer emotions into consideration while strictly protecting personal information, thereby maintaining a competitive advantage in the market.
[0808] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0809] Step 1:
[0810] The user records audio using a smartphone or dedicated device. The input is the user's voice, and the output is an audio data file. The microphone on the device captures the sound and saves it in a file format.
[0811] Step 2:
[0812] The terminal sends recorded audio data to the server. The input is an audio data file, and the output is data transfer to the server. The terminal encrypts the data using a secure protocol and sends it to the specified server address.
[0813] Step 3:
[0814] The server saves the received audio data to a database. The input is the audio data transferred from the terminal, and the output is a file stored in the database. The server verifies the integrity of the data before storing it in the database.
[0815] Step 4:
[0816] The server converts audio data into text information using a speech conversion unit. The input is an audio data file, and the output is text data. Speech recognition software analyzes the audio into text and generates text information.
[0817] Step 5:
[0818] The server uses a personal information identification unit to identify personal information from text data. The input is text data, and the output is the detected personal information. Natural language processing technology is used to extract names, addresses, and other information.
[0819] Step 6:
[0820] The server uses the information protection unit to encrypt the identified personal information. The input is the detected personal information, and the output is the encrypted personal information. An encryption algorithm is applied to ensure security.
[0821] Step 7:
[0822] The server uses an emotion analysis unit to estimate emotional states from voice data. The input is voice data, and the output is a classified emotional state. It analyzes the tone and speed of the voice to estimate emotions such as "joy" or "sadness."
[0823] Step 8:
[0824] The server transmits encrypted information and emotional states to an external service via a data communication unit. The input is encrypted personal information and emotional states, and the output is data transfer to an external service. Data is securely shared using an API.
[0825] (Application Example 2)
[0826] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0827] In today's world, there is a need to accurately monitor emotional changes while ensuring individual safety. In particular, there is a need for a system that can proactively detect and warn of risks associated with rapid emotional shifts. However, existing voice analysis technologies struggle to evaluate a user's emotional state in real time and generate appropriate alerts while maintaining the protection of personal information. Technologies are needed to overcome these limitations.
[0828] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0829] In this invention, the server includes means for acquiring acoustic information, means for converting acoustic information into text information, and means for analyzing the acoustic information to evaluate the user's emotional state. This makes it possible to detect emotional states from acoustic information and generate warnings for abnormal emotional changes while protecting personal information.
[0830] The "data acquisition unit" is a component that has the function of collecting acoustic information from the user's environment.
[0831] The "speech recognition unit" is a component that appropriately converts acquired acoustic information into text information and makes it into an analyzable format.
[0832] The "personal information detection unit" is the part of the system that has the function of identifying information that can identify an individual from textual information and analyzing it securely.
[0833] The "encryption processing unit" is the part of the system that has the function of encrypting identified personal information to protect it from unauthorized access by others.
[0834] The "emotion analysis unit" is a dedicated component that analyzes the tone and rhythm of the user's voice from acoustic information and evaluates their emotional state.
[0835] The "data transmission unit" is a configuration that has the function of transmitting processed data and encrypted information to external functions or services.
[0836] The "warning generation unit" is a component that has the function of issuing an appropriate warning when the user's emotional state exceeds a predetermined threshold.
[0837] To implement this invention, a system is needed that has a series of processes for collecting acoustic information, analyzing that information, and evaluating the user's emotional state. First, acoustic information is collected from the user's surroundings using a terminal such as a smartphone or a dedicated device. This data is securely transmitted to a server through a data acquisition unit. On the server, speech recognition software (e.g., Google Speech-to-Text API) is used to convert the acoustic information into text information.
[0838] The server analyzes textual information using natural language processing technology, and the personal information identified by the personal information detection unit is protected by the encryption processing unit. Meanwhile, the emotion analysis unit uses a generative AI model (e.g., a custom model using TensorFlow) to evaluate the user's emotional state from acoustic information. This process includes the analysis of sound tone and rhythm, and emotions are classified into multiple categories.
[0839] If the emotional state exceeds a set threshold, the server sends an alert to emergency contacts via the warning generation unit. This procedure enables real-time emotion monitoring and warning generation based on acoustic information while maintaining the security of personal information.
[0840] For example, if a user is experiencing significant stress at home, the system immediately detects this change in emotion and sends a warning message stating, "The user may be experiencing stress." This warning is then sent via a smartphone app to pre-configured emergency contacts.
[0841] An example of a prompt message would be: "Use this prompt message to monitor the user's emotional changes. Analyze the emotional information obtained from the voice data and issue appropriate actions according to the urgency." This allows the system to effectively track the user's emotional state and take necessary actions.
[0842] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0843] Step 1:
[0844] The terminal collects acoustic information from the user's surroundings using its built-in microphone. This acoustic information becomes the input data. The terminal then prepares to transmit this data to the server via the data acquisition unit.
[0845] Step 2:
[0846] The server receives the acoustic information transmitted from the terminal in its data acquisition unit. Next, it uses speech recognition software (e.g., Google Speech-to-Text API) to convert this acoustic information into text information. This conversion results in the text information becoming the output data.
[0847] Step 3:
[0848] The server analyzes the text information converted by the speech recognition unit using natural language processing technology, and extracts personally identifiable information in the personal information detection unit. This personal information is the output of the processing and is then securely protected by the encryption processing unit.
[0849] Step 4:
[0850] The server uses an emotion analysis unit to analyze the tone and rhythm of the voice from the acoustic information, and utilizes a generative AI model to evaluate the user's emotional state. In this process, the acoustic information is the input data, and the emotion evaluation result is obtained as output data.
[0851] Step 5:
[0852] The server determines whether the emotional state has exceeded a set threshold based on the emotion evaluation results obtained from the emotion analysis unit. If the threshold is exceeded, the warning generation unit sends an alert to the emergency contact. At this time, the warning alert is output using the emotion evaluation results as input.
[0853] Step 6:
[0854] The user receives an alert from the server and takes the appropriate action according to the instructions in the prompt. In this step, the action based on the prompt appears as output.
[0855] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0856] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0857] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0858] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0859] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0860] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0861] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0862] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0863] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0864] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0865] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0866] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0867] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0868] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0869] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0870] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0871] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0872] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0873] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0874] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0875] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0876] The following is further disclosed regarding the embodiments described above.
[0877] (Claim 1)
[0878] A data receiving unit is provided, along with means for acquiring audio data.
[0879] It includes a speech recognition unit and means for converting the speech data into text data,
[0880] A personal information detection unit is provided, and means are used to identify personal information from the text data,
[0881] Equipped with an encryption processing unit, and means for encrypting identified personal information,
[0882] A data processing system comprising a data transmission unit and means for transmitting the encrypted data to an external service.
[0883] (Claim 2)
[0884] The data processing system according to claim 1, which uses natural language processing technology to identify the aforementioned personal information.
[0885] (Claim 3)
[0886] The data processing system according to claim 1, further comprising a decryption unit for decrypting the encrypted personal information, wherein the decryption unit enables the restoration of personal information in internal processing.
[0887] "Example 1"
[0888] (Claim 1)
[0889] A data acquisition unit and means for storing audio information,
[0890] A voice conversion unit is provided, and means are used to convert the voice information into text information,
[0891] It includes an information detection unit and means for extracting important information from the text information,
[0892] It includes an information protection unit and means for converting extracted important information into data,
[0893] A data transmission unit is provided, and means are used to transmit the converted data to an external platform,
[0894] A means for receiving analysis results from an external platform and performing the necessary processing,
[0895] A system that includes a recovery unit for recovering a portion of the received result data, and includes means for recovering information as needed.
[0896] (Claim 2)
[0897] The system according to claim 1, which uses language processing technology to extract the aforementioned important information.
[0898] (Claim 3)
[0899] The system according to claim 1, which performs analysis on the external platform using a generative algorithm and provides a solution to a specific problem based on the results.
[0900] "Application Example 1"
[0901] (Claim 1)
[0902] It includes a data acquisition unit and means for acquiring voice information,
[0903] A voice conversion unit is provided, and means are used to convert the voice information into text data,
[0904] It includes an information detection unit and means for extracting personal information from the character data,
[0905] A processing unit is provided, and means for encrypting the extracted personal information,
[0906] A transmission unit is provided, and means are used to transmit the encrypted information to an external system,
[0907] A system comprising an analysis unit for feeding back analysis results as internal information, and a means for improving internal response efficiency by utilizing the results processed by the external system.
[0908] (Claim 2)
[0909] The system according to claim 1, which uses natural language processing technology to extract the aforementioned personal information.
[0910] (Claim 3)
[0911] The system according to claim 1, further comprising a decryption unit for decrypting the encrypted personal information, wherein the decryption unit enables the restoration of the personal information in a secure internal environment.
[0912] "Example 2 of combining an emotion engine"
[0913] (Claim 1)
[0914] It includes a data acquisition unit and means for acquiring voice information,
[0915] It includes a voice conversion unit and means for converting the voice information into text information,
[0916] A personal information identification unit is provided, and means are used to identify personal information from the textual information,
[0917] It is equipped with an information protection section and means to encrypt identified personal information,
[0918] It includes an emotion analysis unit and means for estimating emotional states from voice information,
[0919] A system comprising a data communication unit and means for transmitting the encrypted information and emotional state to an external service.
[0920] (Claim 2)
[0921] The system according to claim 1, which uses language analysis technology to identify the aforementioned personal information.
[0922] (Claim 3)
[0923] The system according to claim 1, further comprising a decryption unit for decrypting the encrypted personal information, wherein the decryption unit enables the restoration of the personal information in internal processing.
[0924] "Application example 2 when combining with an emotional engine"
[0925] (Claim 1)
[0926] It includes a data acquisition unit and means for acquiring acoustic information,
[0927] A voice recognition unit is provided, and means are used to convert the acoustic information into text information,
[0928] It includes a personal information detection unit and means for identifying personal information from the text information,
[0929] Equipped with an encryption processing unit, and means for encrypting identified personal information,
[0930] The system includes an emotion analysis unit and means for analyzing the acoustic information to evaluate the user's emotional state,
[0931] It includes a data transmission unit and means for transmitting the encrypted data to an external function,
[0932] A system equipped with a warning generation unit, which includes means for issuing a warning when the user's emotional state exceeds a predetermined threshold.
[0933] (Claim 2)
[0934] The system according to claim 1, which uses natural language processing technology to identify the personal information and a generative AI model to evaluate the emotional state.
[0935] (Claim 3)
[0936] The system according to claim 1, further comprising a decryption unit for decrypting the encrypted personal information, wherein the decryption unit enables the restoration of personal information in internal processing and generates a warning based on the processing result including the sentiment evaluation result. [Explanation of Symbols]
[0937] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A data receiving unit is provided, along with means for acquiring audio data. It includes a speech recognition unit and means for converting the speech data into text data, A personal information detection unit is provided, and means are used to identify personal information from the text data, Equipped with an encryption processing unit, and means for encrypting identified personal information, A data processing system comprising a data transmission unit and means for transmitting the encrypted data to an external service.
2. The data processing system according to claim 1, which uses natural language processing technology to identify the aforementioned personal information.
3. The data processing system according to claim 1, further comprising a decryption unit for decrypting the encrypted personal information, wherein the decryption unit enables the restoration of personal information in internal processing.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A