system

A voice-activated system with audio and video guides, multilingual support, and on-site assistance addresses the digital divide by offering personalized and efficient technical support.

JP2026070874APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The digital divide remains a significant issue, particularly affecting the elderly and those with low digital literacy, as they face challenges in accessing effective technical support due to delays and unclear explanations, especially in regions with underdeveloped infrastructure.

Method used

A system that receives voice input, performs voice recognition, and provides operation guides through audio and video, with multilingual support, customizable to the user's skill level, and offers on-site assistance when needed.

Benefits of technology

Effectively bridges the digital divide by providing tailored, accessible, and comprehensive support to users with diverse language and skill levels, ensuring prompt and clear assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070874000001_ABST
    Figure 2026070874000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A terminal that receives voice commands from the user, A server means that performs speech recognition in order to analyze audio data, A server means that selects and provides an operation guide based on the analysis results, A terminal device that presents the selected operation guide to the user via audio and video, A terminal device that receives user feedback and sends it to a server, A server mechanism to notify local support staff and arrange for on-site assistance if the user's problem cannot be resolved. Includes system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern information society, the digital divide remains a prominent problem, where the elderly and individuals with low digital literacy cannot fully enjoy the benefits of technology. Such a situation tends to be more serious in areas where regional infrastructure and support systems are not well-developed. Current support is prone to problems such as delays in inquiry responses and unclear technical explanations, and is insufficient to quickly and effectively solve users' technical problems.

Means for Solving the Problems

[0005] This invention provides a system that effectively bridges the digital divide by having a terminal receive voice input from a user, performing voice recognition on a server, and providing appropriate operation guides based on the analysis results. Furthermore, the operation guides are presented via both audio and video, and multilingual support ensures that users with different language settings receive appropriate language support. In addition, the system is customized according to the user's technical skill level, and if a problem remains unresolved, on-site assistance by local support staff is arranged, enabling comprehensive support.

[0006] A "user" is an individual or group that uses the system to issue voice commands.

[0007] A "voice command" is a voice command used by a user to communicate operational instructions to a system.

[0008] A "terminal" is a device that has the function of receiving voice input from a user and sending it to a server.

[0009] A "server" is a central device that analyzes received audio data, selects the appropriate operation guide, and transmits it to the terminal.

[0010] "Speech recognition" is a technology that analyzes received audio data and converts it into text information.

[0011] An "operation guide" is a set of instructions that shows the operating procedures for a digital device according to the user's requirements.

[0012] "Presented with audio and video" means that in addition to audio instructions, visual demonstrations are displayed on the screen.

[0013] "Multilingual support" is a function that allows information to be provided in the language of the user, regardless of their language.

[0014] "Technical skill level" is a measure indicating the user's level of understanding and operation ability regarding digital devices and technologies.

[0015] "Regional support staff" refers to professional workers who provide direct technical support in the user's region.

[0016] "On-site support" is an activity in which regional support staff visit the user's location to provide direct support.

Brief Explanation of Drawings

[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] The system of this invention supports bridging the digital divide by allowing users to give voice commands for operation. Users use terminals such as smartphones or computers to inquire by voice about operations that require support.

[0039] After receiving a voice command from the user, the terminal sends the data to the server via the internet. The server converts the received voice data into text using a speech recognition engine. Next, it uses a natural language processing (NLP) engine to analyze the user's request and identify what kind of support is needed.

[0040] Based on the analysis results, the server selects the appropriate guide from its stored operation guides. These guides include audio and video instructions and are presented in a highly user-friendly format. The selected guide is tailored to the user's skill level and language settings.

[0041] Audio and video content generated by the server is sent to the terminal and provided to the user. The terminal supports the user in easily understanding and performing operations by displaying visual demonstrations on the screen along with voice instructions. For example, if the user asks by voice, "Tell me how to send an email," the server provides a series of video guides from basic settings in the email application to creating and sending an email.

[0042] When a user provides feedback, the device sends that information back to the server. The server analyzes the feedback, evaluates whether the user's problem has been resolved, and provides additional support as needed. If the user's problem is technically complex and cannot be resolved, the server can notify local support staff and schedule an on-site assistance visit.

[0043] This system features multilingual support and customization options tailored to different skill levels, enabling it to provide equally accurate support to users with diverse backgrounds. Its aim is to assist individuals who feel uneasy about using digital devices and effectively bridge the digital divide.

[0044] The following describes the processing flow.

[0045] Step 1:

[0046] The user issues voice commands to a device such as a smartphone or PC. The commands are recorded by the microphone built into the device.

[0047] Step 2:

[0048] The device sends the recorded audio data to the server via the internet. This communication is encrypted and conducted in a secure manner.

[0049] Step 3:

[0050] The server processes the received audio data through a speech recognition engine, converting the speech into text. Speech recognition is performed with high accuracy using an AI model.

[0051] Step 4:

[0052] The server analyzes the converted text using a natural language processing (NLP) engine to identify the type of support the user is seeking. This analysis helps understand the user's intent and the context of their question.

[0053] Step 5:

[0054] Based on the analysis results, the server selects the appropriate operation guide from the database. This guide is customized according to the user's technical skill level and language settings.

[0055] Step 6:

[0056] The server generates audio and video instructions based on the selected operation guide and sends them to the terminal. These instructions are presented in a format that is easy for the user to understand.

[0057] Step 7:

[0058] The device displays audio and video instructions received by the user. In addition to audio explanations, visual demonstrations are provided on the screen.

[0059] Step 8:

[0060] The system can perform actions based on the information provided by the user and provide feedback on whether the support was helpful. This feedback may also be provided via voice.

[0061] Step 9:

[0062] The device sends user feedback back to the server. Based on this feedback, the server determines whether to improve support or if additional support is needed.

[0063] Step 10:

[0064] If the user's problem is not resolved, the server will notify local support staff. On-site support will be arranged depending on the situation.

[0065] (Example 1)

[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0067] With the evolution of digital technology, a problem arises where users with limited technical skills and knowledge are unable to effectively utilize digital devices and services. This digital divide is particularly pronounced in multilingual environments and among individuals with diverse technical skill sets, hindering access to information. Furthermore, there is the challenge of difficulty in receiving prompt and appropriate support when technical assistance is needed.

[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0069] In this invention, the server includes a computing device means for performing automatic speech recognition to convert speech data into text information, a language analysis means for analyzing the text information and identifying the requested support, and a computing device means for selecting and providing appropriate operating guidelines based on the analysis results. This enables users to easily operate digital devices and services and receive necessary support in multiple languages ​​and adaptively, regardless of their technical skills or knowledge.

[0070] "Computer means" refers to a device that has the function of receiving voice instructions from the user and presenting selected operating guidelines in both audio and video.

[0071] The "computational device means" is a device that performs automatic speech recognition to convert speech data into text information, and further analyzes the text information to identify the requested support content.

[0072] "Language analysis means" refers to technology that analyzes text data converted by speech recognition to understand user requests.

[0073] "Operating guidelines" are audio and video content that explains the procedures for specific operations or tasks, providing effective guidance to users.

[0074] A "technical support provider" is a specialist or staff member who provides assistance when technical problems cannot be resolved.

[0075] "On-site support" refers to technical support personnel traveling to the user's location to provide direct assistance.

[0076] "Multilingual processing" is a technology that has the capability to provide appropriate support in each language to users who use different languages.

[0077] "Adaptation" refers to the process by which a system adjusts itself to provide optimal support based on the user's technical skill level.

[0078] The system of the present invention aims to bridge the digital divide by enabling users to request the operation of digital devices using voice commands. Users input voice commands using terminal devices such as smartphones or computers, and this voice data is transmitted to a server.

[0079] The server uses a speech recognition engine to process the received audio data. This engine can be a commonly used speech recognition software, such as Google® Speech-to-Text API. This converts the audio data into text information.

[0080] Next, the server uses a natural language processing (NLP) engine to analyze the transcribed data. This allows it to accurately understand the intent of the user's request and identify the necessary assistance. Common natural language processing tools can be used for this analysis.

[0081] Once the analysis is complete, the server selects the appropriate guide from several operation guides stored in storage and configures its contents. This guide includes detailed audio and visual instructions, making it easy for the user to understand and perform the operation. Customization is also possible according to the user's technical skill level and language settings, and the user interface is adjusted based on the language and skills selected by the user.

[0082] The device provides the user with audio and video guides received from the server. This device displays visual demonstrations and, in conjunction with audio instructions, assists the user in smooth operation. For example, if the user asks "How to create a homepage," the server provides a beginner-friendly HTML / CSS basics guide and plays a visualized version of that guide on the device.

[0083] Furthermore, when a user submits feedback, this feedback information is sent back to the server, which analyzes it and evaluates whether the user's problem has been resolved. If the problem is difficult to resolve, the server notifies local technical support staff, and on-site assistance is arranged.

[0084] Example prompt: "Create a beginner's guide to HTML / CSS basics based on the user's voice command 'How to create a homepage'."

[0085] This allows users to efficiently utilize digital devices regardless of their technical knowledge.

[0086] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0087] Step 1:

[0088] The user inputs voice commands using the device. Specifically, they use the device's microphone to issue voice commands such as "I want to edit the photo." The input is the user's voice command, and the output is audio data.

[0089] Step 2:

[0090] The terminal sends the acquired audio data directly to the server. The input is the user's audio data, and the output is the audio data that the server uses to begin processing. The transmission process involves data transfer from the terminal to the server.

[0091] Step 3:

[0092] The server converts the received audio data into text information using a speech recognition engine. In this process, the input is audio data and the output is text data. The speech recognition engine analyzes the audio data and generates the corresponding text.

[0093] Step 4:

[0094] The server uses the converted text data to perform natural language processing (NLP) and analyze the user's instructions. The input is text data, and the output is the analysis result. The NLP engine understands the intent of the text and identifies the necessary support.

[0095] Step 5:

[0096] The server selects an operation guide based on the analysis results. The input is the analysis results, and the output is the selected operation guide. The server selects the most appropriate guide from its internal database.

[0097] Step 6:

[0098] The server customizes the selected operation guide to match the user's skill level and language. The input is the operation guide and user configuration information, and the output is the customized guide. The server adjusts the guide content to suit the user.

[0099] Step 7:

[0100] The server transmits customized audio and video content to the terminal. The input is the adjusted guide content, and the output is the content available to the user. The content is delivered to the user's terminal via data communication.

[0101] Step 8:

[0102] The terminal provides the user with received audio and video guides and plays visual demonstrations on the screen. The input is guide content from the server, and the output is displayed content that facilitates user understanding. The terminal provides an environment where it can view the content, supporting the user experience.

[0103] Step 9:

[0104] The user views the guide provided on the terminal, performs the operations, and sends feedback to the terminal as needed. The input is the user's operational feedback, and the output is feedback data. This information is then sent back to the server for further analysis.

[0105] (Application Example 1)

[0106] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0107] There is a need to overcome the difficulties faced by individuals unfamiliar with digital technology when efficiently using safety devices and information processing systems in homes and workplaces, and to provide support tailored to diverse language and technical skill levels. The goal is to enable accurate and rapid responses in emergencies and to allow for the safe use of equipment.

[0108] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0109] In this invention, the server includes means including an information terminal device for receiving voice commands from a user; computer means for performing voice recognition processing for analyzing voice data; computer means for selecting and providing instructions based on the analysis results; means including an information terminal device for presenting the selected instructions to the user audibly and visually; means including an information terminal device for receiving user feedback and transmitting it to the computer; computer means for notifying local support staff and arranging on-site support if the user's problem is not resolved; means for managing safety protection devices based on voice operation; and means for providing voice guidance in emergencies. This enables even users unfamiliar with digital technology to intuitively and quickly operate safety protection devices and other information systems.

[0110] An "information terminal device" is a device that receives voice commands from a user and processes them in an appropriate manner.

[0111] "Computing means" refers to a computing device that has functions for processing digital data, such as speech recognition and analysis of instructions.

[0112] "Speech recognition processing" is the process of analyzing audio data received from a user and converting it into text data.

[0113] A "safety protection device" is a device that has security functions to protect people or assets.

[0114] "Means of providing voice guidance in emergencies" refers to a function that provides voice guidance to encourage users to take appropriate action when an emergency occurs.

[0115] A "local support provider" refers to a support provider who can visit users in person to provide assistance in order to solve their technical problems.

[0116] "User feedback" refers to the process of users providing results and opinions on their experiences, and includes the system receiving and analyzing this feedback.

[0117] "Instructions" refers to operation guides and audio / video instructions provided based on user requests or questions.

[0118] The embodiments for carrying out the invention are described below.

[0119] This invention provides a system that allows users to manage and operate devices by inputting voice commands through an information terminal device and using voice recognition. First, the user sends a voice command from the information terminal device. The terminal collects this command as voice data and transmits it to a computing means. The computing means, in particular, voice recognition software such as the Google Cloud Speech-to-Text API, converts this into text data and further analyzes the content of the instruction using the Google Cloud Natural Language API.

[0120] Based on the analysis results, appropriate instructions are sent back to the information terminal device. The user receives the instructions via voice and visual interfaces and, based on them, can, for example, remotely control the settings of a home safety device by voice. User feedback is also sent back from the terminal to the computer, and arrangements are made to notify local support staff if the user's problem is not resolved.

[0121] This system includes a concrete example where the settings of a security camera are automatically adjusted when the user inputs a prompt as a voice command, such as "I want to be notified when my child comes home." The aim is to bridge the digital divide when managing and operating safety devices through such voice commands.

[0122] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0123] Step 1:

[0124] The user inputs voice commands into the information terminal device. At this time, the terminal uses its built-in microphone to collect voice data and saves that voice data locally as a binary file.

[0125] Step 2:

[0126] The device sends the saved audio data to the server. The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. This process uses acoustic and language models to transform the audio signal into text. Text data is then generated.

[0127] Step 3:

[0128] The server passes the generated text data to the Google Cloud Natural Language API. The natural language processing engine analyzes the text and identifies the intent of the prompt. This process extracts keywords from the text and analyzes the context to understand the user's request. The intent is obtained as a result of the analysis.

[0129] Step 4:

[0130] The server consults a database to select appropriate instructions based on the user's intent. This database contains pre-prepared audio and visual guide templates. The server searches for and selects instruction content that matches the user's request. The instructions are then finalized as output.

[0131] Step 5:

[0132] The server transmits the selected instructions to the terminal as audio and visual media. The terminal plays these instructions and displays a visual demonstration on the screen. The user can perform the task according to the instructions by listening to the audio instructions and referring to the visual guide.

[0133] Step 6:

[0134] After completing an operation, the user enters feedback through an information terminal device. The terminal sends this feedback data to a server. The server analyzes the feedback and evaluates whether the problem has been resolved. This analysis involves determining the content of the feedback through text analysis and notifying local support personnel if additional support is needed.

[0135] This series of steps allows users to intuitively and efficiently manage and operate equipment using voice commands.

[0136] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0137] In the system of this invention, not only does the user instruct operations using voice commands, but the system also estimates the user's emotional state by analyzing the voice. This provides more personalized support and improves the user experience.

[0138] Users request support via devices such as smartphones or computers. After the device receives the user's voice command, it sends the voice data to the server. At this time, the server, equipped with an emotion engine, analyzes the voice data and processes it to identify the user's emotions. The emotion engine determines the user's emotions based on characteristics such as the intonation, speed, and pitch of the voice.

[0139] Based on the results of the emotion analysis, the server takes into account the user's emotional state and selects the most appropriate operation guide. For example, if the emotion engine determines that the user is confused, it aims to alleviate the user's anxiety by selecting a guide that provides detailed step-by-step instructions.

[0140] The user guide is tailored to the user's technical skill level, language, and emotional state, and is sent from the server to the terminal. This ensures the guide is tailored to the user's needs, enabling more effective support. Users can follow the instructions while viewing the audio and video displayed on the terminal.

[0141] Furthermore, users can provide feedback on the usefulness of the support, and this feedback is resent to the server via their device. The server analyzes the feedback and, if necessary, compares it with the results of the emotion engine to improve future support methods. Also, if the emotion engine determines from the feedback that the user's emotions remain unstable, it notifies local support staff to prioritize on-site assistance.

[0142] In this way, by utilizing an emotion engine, this system aims to not only solve technical problems but also comprehensively bridge the digital divide by taking into account the psychological aspects of the user.

[0143] The following describes the processing flow.

[0144] Step 1:

[0145] The user inputs a voice command into the device. The device records this using its microphone and temporarily stores it as audio data.

[0146] Step 2:

[0147] The device sends the recorded audio data to the server. This transmission is encrypted to prevent eavesdropping on the audio data.

[0148] Step 3:

[0149] The server receives the audio data and uses a speech recognition engine to convert the audio data into text.

[0150] Step 4:

[0151] The server analyzes the converted text using a natural language processing (NLP) engine to identify the type of support the user is requesting.

[0152] Step 5:

[0153] The server processes the voice data into an emotion engine, analyzing factors such as tone, intensity, and speed to estimate the user's emotional state.

[0154] Step 6:

[0155] The server combines the results of natural language processing and sentiment analysis to select the most appropriate operation guide from the database. This selection takes into account the user's pre-configured technical skill level and emotional state.

[0156] Step 7:

[0157] The server converts the selected operation guide into audio and video instructions and sends them to the terminal. These instructions are designed to be easy for the user to understand.

[0158] Step 8:

[0159] The device provides audio and video instructions to the user. The user can proceed with the operation while watching the on-screen demonstration.

[0160] Step 9:

[0161] Users can provide voice feedback. The device records that feedback and sends it to the server.

[0162] Step 10:

[0163] The server receives feedback and analyzes it using the emotion engine. If the user's emotions remain unstable, it notifies local support staff and prioritizes arranging on-site assistance.

[0164] (Example 2)

[0165] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0166] As modern information and communication technologies advance, users are increasingly required to operate devices based on diverse voice commands. However, conventional systems only perform simple voice recognition without considering the user's emotional state, making it difficult to provide effective support. Furthermore, they lack the flexibility to adapt to the diversity of users' technical abilities and languages. This can potentially expose users to stress and anxiety, potentially impairing their digital technology experience.

[0167] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0168] In this invention, the server includes a computing device means that performs speech recognition and emotion estimation to analyze voice information, a computing device means that selects and presents operational guidance based on the analysis results, and a computing device means that notifies local support personnel and arranges on-site support if the user's problem is not resolved. This enables the provision of personalized operational guides that correspond to the user's emotional state and flexible support that accommodates diversity in technical abilities and languages.

[0169] A "user" refers to an individual who operates the system and issues voice commands.

[0170] "Voice commands" refer to information that users give via voice commands or requests through their devices.

[0171] "Device means" refers to hardware or software functions for receiving voice commands from the user or displaying information.

[0172] "Voice information" refers to the voice data obtained after converting the user's voice commands into digital data.

[0173] "Speech recognition" refers to the technology that analyzes speech information and understands the content of speech commands.

[0174] "Emotion estimation" refers to a technology that determines a user's emotional state by analyzing voice information.

[0175] "Computation device means" refers to a computer system for speech recognition, emotion estimation, and the selection and presentation of operational guidance.

[0176] "Operational guidance" refers to information that provides guidelines and procedures for taking action based on the user's voice commands.

[0177] An "information display device" refers to hardware such as displays and projection equipment used to provide users with instructions on how to operate a system.

[0178] An "image device" refers to a device used to visually present information to a user.

[0179] A "regional support officer" refers to a specialist designated to provide on-site assistance in resolving users' technical problems.

[0180] "On-site support" refers to support activities that involve actually going to the user's location and directly resolving their problems.

[0181] "Multilingual support" refers to a feature that provides assistance in the language each user understands, for users with different language settings.

[0182] "Technical ability" refers to the technical knowledge and operational skills that a user possesses.

[0183] This invention is a system that improves the user experience by allowing users to operate it through voice commands. Users use a terminal such as a smartphone or computer to issue voice commands. The terminal converts the voice into digital data using a microphone or voice input module and transmits it to a server for analysis.

[0184] The server functions as a computing device equipped with advanced emotion estimation capabilities. This server analyzes audio data using speech recognition algorithms to estimate the user's emotional state. It analyzes features such as intonation, speed, and pitch, and uses generative AI models to determine the emotional state. For example, it constructs emotion estimation models using machine learning frameworks such as TENSORFLOW® or PyTorch.

[0185] Based on the analysis results, the server selects appropriate instructions. These instructions are customized based on the user's emotional state, technical ability, and multilingual support. The server-selected instructions are sent to the terminal, which then provides guidance to the user using a video display or audio output device.

[0186] As a concrete example, suppose a user gives a voice command saying, "I want to set up email on my new device." In this case, the server will assume the user is feeling anxious and select a detailed, step-by-step email setup guide for first-time users, sending it to the device. Alternatively, it can send a prompt to the generative AI model such as: "The user is anxious about setting up email on their new device. Please suggest a gentle and detailed email setup guide." This allows the user to proceed with confidence.

[0187] Furthermore, the server collects user feedback and evaluates the usefulness of voice guidance. Feedback analysis allows for continuous improvement of support services and responses to user requests. By notifying local support staff and arranging on-site assistance as needed, more comprehensive support becomes possible.

[0188] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0189] Step 1:

[0190] The user issues voice commands to the terminal. The input is the user's voice. The terminal uses its microphone to receive the voice as an analog signal. A voice processing module converts the analog signal into digital data, performing noise reduction and sound quality improvement. This digital data is output in the form of voice commands and is ready to be sent to the server.

[0191] Step 2:

[0192] The terminal sends digitized voice commands to the server via a security protocol. The input is digital voice data. The terminal performs data transfer, and the server receives the voice data. The output is the digital voice data transferred to the server.

[0193] Step 3:

[0194] The server analyzes the received digital audio data. The input is audio data. The server uses a speech recognition engine to convert it into text-based voice commands, and then utilizes an emotion estimation engine to determine the user's emotions. It analyzes data such as intonation, speed, and pitch, and outputs an emotion index.

[0195] Step 4:

[0196] The server selects the optimal operation guidance based on voice commands and emotion indicators. The inputs are voice commands and emotion indicators. Using a generative AI model, it infers the most appropriate operation guidance from the voice commands and further adjusts the guidance considering the emotion indicators. The output is customized operation guidance.

[0197] Step 5:

[0198] The server sends the configured operating instructions to the terminal. These instructions include video explanations and audio guides. The input is the customized operating instructions, which are sent to the terminal via data transfer from the server. The output is the operating instructions received by the terminal.

[0199] Step 6:

[0200] The terminal displays operating instructions received from the server to the user. The input is the operating instruction data. The terminal's display and speaker are used to present information to the user visually and audibly. The output is the operating instructions displayed to the user.

[0201] Step 7:

[0202] The user inputs feedback on the usefulness of the operation instructions into the terminal. This input is user feedback information. The terminal collects the input feedback and sends it to the server. The output of this transmission is the feedback data received by the server.

[0203] Step 8:

[0204] The server analyzes the received feedback and considers ways to improve future support. The input is feedback data. It compares this data with sentiment estimation results to evaluate system performance and identify areas for improvement. Furthermore, it sends notifications to local support staff as needed. The output is an improved support plan or support notification.

[0205] (Application Example 2)

[0206] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0207] Current content distribution services are insufficient in suggesting content that is tailored to the individual emotional state of users. Therefore, there is a need for services that are more attentive to users' emotions.

[0208] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0209] In this invention, the server includes means for receiving voice commands from the user and analyzing the voice data to estimate the user's emotional state; means for selecting an operation guide based on the results of voice recognition and emotion analysis and optimizing it according to the user's emotional state; and means for presenting the selected operation guide to the user visually and audibly. This enables personalized content suggestions based on the user's emotional state.

[0210] "Voice commands" refer to a method in which users give specific operational instructions to a system using their voice.

[0211] "Voice data" refers to the digital representation of voice input collected from users.

[0212] "Emotional state" refers to the emotional tendencies and psychological state analyzed from the user's voice.

[0213] "Speech recognition" is a technology that analyzes input speech data and converts it into text.

[0214] An "operation guide" is a set of instructions and procedures provided to a user to operate a system.

[0215] "Presenting information visually and aurally" refers to methods of conveying information to users using screen displays and audio output.

[0216] "Feedback" refers to the act of a user responding to or evaluating the services provided by a system.

[0217] "Emotional analysis" is a process that analyzes voice data to identify the user's emotional state.

[0218] "Improving support services" means optimizing the services and guides provided based on user feedback.

[0219] This invention is a personalized content delivery system that takes into account the user's emotional state, and primarily uses the user's voice input to analyze their emotions and provide optimal content. The system is implemented using the following means.

[0220] The system's terminal receives voice commands from the user. The voice data is sent to the server via the internet. The server converts the voice data into text using the Google Cloud Speech-to-Text API. Subsequently, it uses Azure® Text Analytics to perform sentiment analysis on this text data and estimate the user's emotional state.

[0221] Once the user's emotional state has been estimated, the server selects operation guides and content based on the results. The selection criteria include considering the user's psychological feedback and requests such as wanting to relax or become excited. During this process, generative AI models are used to optimize the user interface and operation experience.

[0222] Finally, the selected content and guides are presented to the device visually and aurally. Users can interact with these instructions and provide feedback. This feedback information is then sent back to the server to help improve the quality of future content delivery.

[0223] As a concrete example, consider a scenario where a user watches a movie and says, "Tell me what movie you recommend I watch next." If sentiment analysis of the voice data suggests the user is in a happy mood, a popular local comedy film will be suggested based on this. An example of a prompt would be, "Analyze the voice feedback the user has given on the content they watched and recommend content based on their emotional state."

[0224] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0225] Step 1:

[0226] The terminal receives voice commands from the user. Voice data is input and digitized as an audio file. The received voice data is transferred to a server via the intranet for processing.

[0227] Step 2:

[0228] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. This process transforms the audio data into a parseable text format, making it ready for use in subsequent processes.

[0229] Step 3:

[0230] The server uses Azure Text Analytics to perform sentiment analysis on the converted text data. Here, text data is input, and evaluation metrics indicating the user's emotional state are generated. This quantifies and outputs the user's psychological state.

[0231] Step 4:

[0232] The server selects appropriate content or user guides based on the sentiment analysis results. Here, sentiment evaluation metrics and the user's past history data are used as input, and a generative AI model is employed to select the optimal content. This output forms the basis for the next step.

[0233] Step 5:

[0234] The selected content or operating guide is sent to the device. The content is entered into the device and presented to the user visually and aurally. Through this process, the user receives personalized content.

[0235] Step 6:

[0236] The user provides voice feedback on the content presented by the system. The device then records the user's feedback again as voice data and prepares to send it to the server.

[0237] Step 7:

[0238] The server receives feedback from the user and performs sentiment analysis again. Using the feedback audio data as input, it outputs a new sentiment evaluation metric. This allows the server to confirm user satisfaction and use the results to improve service quality.

[0239] Step 8:

[0240] Based on feedback, the server improves its content selection algorithm and sentiment analysis model as needed. This improvement process uses accumulated sentiment data and feedback as input, enabling the delivery of more accurate content.

[0241] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0242] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0243] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0244] [Second Embodiment]

[0245] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0246] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0247] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0248] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0249] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0250] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0251] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0252] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0253] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0254] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0255] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0256] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0257] The system of this invention supports bridging the digital divide by allowing users to give voice commands for operation. Users use terminals such as smartphones or computers to inquire by voice about operations that require support.

[0258] After receiving a voice command from the user, the terminal sends the data to the server via the internet. The server converts the received voice data into text using a speech recognition engine. Next, it uses a natural language processing (NLP) engine to analyze the user's request and identify what kind of support is needed.

[0259] Based on the analysis results, the server selects the appropriate guide from its stored operation guides. These guides include audio and video instructions and are presented in a highly user-friendly format. The selected guide is tailored to the user's skill level and language settings.

[0260] Audio and video content generated by the server is sent to the terminal and provided to the user. The terminal supports the user in easily understanding and performing operations by displaying visual demonstrations on the screen along with voice instructions. For example, if the user asks by voice, "Tell me how to send an email," the server provides a series of video guides from basic settings in the email application to creating and sending an email.

[0261] When a user provides feedback, the device sends that information back to the server. The server analyzes the feedback, evaluates whether the user's problem has been resolved, and provides additional support as needed. If the user's problem is technically complex and cannot be resolved, the server can notify local support staff and schedule an on-site assistance visit.

[0262] This system features multilingual support and customization options tailored to different skill levels, enabling it to provide equally accurate support to users with diverse backgrounds. Its aim is to assist individuals who feel uneasy about using digital devices and effectively bridge the digital divide.

[0263] The following describes the processing flow.

[0264] Step 1:

[0265] The user issues voice commands to a device such as a smartphone or PC. The commands are recorded by the microphone built into the device.

[0266] Step 2:

[0267] The device sends the recorded audio data to the server via the internet. This communication is encrypted and conducted in a secure manner.

[0268] Step 3:

[0269] The server processes the received audio data through a speech recognition engine, converting the speech into text. Speech recognition is performed with high accuracy using an AI model.

[0270] Step 4:

[0271] The server analyzes the converted text using a natural language processing (NLP) engine to identify the type of support the user is seeking. This analysis helps understand the user's intent and the context of their question.

[0272] Step 5:

[0273] Based on the analysis results, the server selects the appropriate operation guide from the database. This guide is customized according to the user's technical skill level and language settings.

[0274] Step 6:

[0275] The server generates audio and video instructions based on the selected operation guide and sends them to the terminal. These instructions are presented in a format that is easy for the user to understand.

[0276] Step 7:

[0277] The device displays audio and video instructions received by the user. In addition to audio explanations, visual demonstrations are provided on the screen.

[0278] Step 8:

[0279] Operations can be performed based on the information provided by the user, and the terminal can provide feedback on whether the support was helpful. This feedback may be provided audibly.

[0280] Step 9:

[0281] The terminal resends the feedback from the user to the server. Based on this feedback, the server determines whether to improve the support content or the need for additional support.

[0282] Step 10:

[0283] If the user's problem is not solved, the server notifies the regional support staff. On-site support is arranged according to the situation.

[0284] (Example 1)

[0285] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0286] With the evolution of digital technology, there is a problem that users with limited technical skills and knowledge cannot effectively use digital devices and services. This digital divide is particularly prominent among individuals in multilingual environments or with different technical skills, becoming an obstacle to information access. There is also an issue that it is difficult to receive prompt and appropriate support in situations where technical support is required.

[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0288] In this invention, the server includes a computing device means for performing automatic speech recognition to convert speech data into text information, a language analysis means for analyzing the text information and identifying the requested support, and a computing device means for selecting and providing appropriate operating guidelines based on the analysis results. This enables users to easily operate digital devices and services and receive necessary support in multiple languages ​​and adaptively, regardless of their technical skills or knowledge.

[0289] "Computer means" refers to a device that has the function of receiving voice instructions from the user and presenting selected operating guidelines in both audio and video.

[0290] The "computational device means" is a device that performs automatic speech recognition to convert speech data into text information, and further analyzes the text information to identify the requested support content.

[0291] "Language analysis means" refers to technology that analyzes text data converted by speech recognition to understand user requests.

[0292] "Operating guidelines" are audio and video content that explains the procedures for specific operations or tasks, providing effective guidance to users.

[0293] A "technical support provider" is a specialist or staff member who provides assistance when technical problems cannot be resolved.

[0294] "On-site support" refers to technical support personnel traveling to the user's location to provide direct assistance.

[0295] "Multilingual processing" is a technology that has the capability to provide appropriate support in each language to users who use different languages.

[0296] "Adaptation" refers to the process by which a system adjusts itself to provide optimal support based on the user's technical skill level.

[0297] The system of the present invention aims to bridge the digital divide by enabling users to request the operation of digital devices using voice commands. Users input voice commands using terminal devices such as smartphones or computers, and this voice data is transmitted to a server.

[0298] The server uses a speech recognition engine to process the received audio data. This engine can be a commonly used speech recognition software, such as the Google Speech-to-Text API. This converts the audio data into text information.

[0299] Next, the server uses a natural language processing (NLP) engine to analyze the transcribed data. This allows it to accurately understand the intent of the user's request and identify the necessary assistance. Common natural language processing tools can be used for this analysis.

[0300] Once the analysis is complete, the server selects the appropriate guide from several operation guides stored in storage and configures its contents. This guide includes detailed audio and visual instructions, making it easy for the user to understand and perform the operation. Customization is also possible according to the user's technical skill level and language settings, and the user interface is adjusted based on the language and skills selected by the user.

[0301] The device provides the user with audio and video guides received from the server. This device displays visual demonstrations and, in conjunction with audio instructions, assists the user in smooth operation. For example, if the user asks "How to create a homepage," the server provides a beginner-friendly HTML / CSS basics guide and plays a visualized version of that guide on the device.

[0302] Also, when the user sends feedback, this feedback information is sent back to the server again. The server performs analysis and evaluates whether the user's problem has been solved. If it is difficult to solve, it is notified to the local technical support staff by the server, and the on-site support is adjusted.

[0303] Example of a prompt sentence: "Based on the user's voice command 'How to create a homepage', please create a basic guide to HTML / CSS for beginners."

[0304] As a result, regardless of the presence or absence of technical knowledge, users can efficiently use digital devices.

[0305] The flow of the specific process in Example 1 will be described using FIG. 11.

[0306] Step 1:

[0307] The user uses the terminal to input a voice instruction. Specifically, using the microphone of the terminal, a voice command such as "I want to edit a photo" is issued. The input is the user's voice instruction, and the output is voice data.

[0308] Step 2:

[0309] The terminal directly sends the acquired voice data to the server. The input is the user's voice data, and the output is the voice data for the server to start processing. Through the transmission, data transfer from the terminal to the server is performed.

[0310] Step 3:

[0311] The server converts the received voice data into character information using a voice recognition engine. In this process, the input is voice data, and the output is data in text format. The voice recognition engine analyzes the voice data and generates the corresponding text.

[0312] Step 4:

[0313] The server uses the converted text data to perform natural language processing (NLP) and analyze the user's instructions. The input is text data, and the output is the analysis result. The NLP engine understands the intent of the text and identifies the necessary support.

[0314] Step 5:

[0315] The server selects an operation guide based on the analysis results. The input is the analysis results, and the output is the selected operation guide. The server selects the most appropriate guide from its internal database.

[0316] Step 6:

[0317] The server customizes the selected operation guide to match the user's skill level and language. The input is the operation guide and user configuration information, and the output is the customized guide. The server adjusts the guide content to suit the user.

[0318] Step 7:

[0319] The server transmits customized audio and video content to the terminal. The input is the adjusted guide content, and the output is the content available to the user. The content is delivered to the user's terminal via data communication.

[0320] Step 8:

[0321] The terminal provides the user with received audio and video guides and plays visual demonstrations on the screen. The input is guide content from the server, and the output is displayed content that facilitates user understanding. The terminal provides an environment where it can view the content, supporting the user experience.

[0322] Step 9:

[0323] The user views the guide provided on the terminal, performs the operations, and sends feedback to the terminal as needed. The input is the user's operational feedback, and the output is feedback data. This information is then sent back to the server for further analysis.

[0324] (Application Example 1)

[0325] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0326] There is a need to overcome the difficulties faced by individuals unfamiliar with digital technology when efficiently using safety devices and information processing systems in homes and workplaces, and to provide support tailored to diverse language and technical skill levels. The goal is to enable accurate and rapid responses in emergencies and to allow for the safe use of equipment.

[0327] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0328] In this invention, the server includes means including an information terminal device for receiving voice commands from a user; computer means for performing voice recognition processing for analyzing voice data; computer means for selecting and providing instructions based on the analysis results; means including an information terminal device for presenting the selected instructions to the user audibly and visually; means including an information terminal device for receiving user feedback and transmitting it to the computer; computer means for notifying local support staff and arranging on-site support if the user's problem is not resolved; means for managing safety protection devices based on voice operation; and means for providing voice guidance in emergencies. This enables even users unfamiliar with digital technology to intuitively and quickly operate safety protection devices and other information systems.

[0329] An "information terminal device" is a device that receives voice commands from a user and processes them in an appropriate manner.

[0330] "Computing means" refers to a computing device that has functions for processing digital data, such as speech recognition and analysis of instructions.

[0331] "Speech recognition processing" is the process of analyzing audio data received from a user and converting it into text data.

[0332] A "safety protection device" is a device that has security functions to protect people or assets.

[0333] "Means of providing voice guidance in emergencies" refers to a function that provides voice guidance to encourage users to take appropriate action when an emergency occurs.

[0334] A "local support provider" refers to a support provider who can visit users in person to provide assistance in order to solve their technical problems.

[0335] "User feedback" refers to the process of users providing results and opinions on their experiences, and includes the system receiving and analyzing this feedback.

[0336] "Instructions" refers to operation guides and audio / video instructions provided based on user requests or questions.

[0337] The embodiments for carrying out the invention are described below.

[0338] This invention provides a system that allows users to manage and operate devices by inputting voice commands through an information terminal device and using voice recognition. First, the user sends a voice command from the information terminal device. The terminal collects this command as voice data and transmits it to a computing means. The computing means, in particular, voice recognition software such as the Google Cloud Speech-to-Text API, converts this into text data and further analyzes the content of the instruction using the Google Cloud Natural Language API.

[0339] Based on the analysis results, appropriate instructions are sent back to the information terminal device. The user receives the instructions via voice and visual interfaces and, based on them, can, for example, remotely control the settings of a home safety device by voice. User feedback is also sent back from the terminal to the computer, and arrangements are made to notify local support staff if the user's problem is not resolved.

[0340] This system includes a concrete example where the settings of a security camera are automatically adjusted when the user inputs a prompt as a voice command, such as "I want to be notified when my child comes home." The aim is to bridge the digital divide when managing and operating safety devices through such voice commands.

[0341] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0342] Step 1:

[0343] The user inputs voice commands into the information terminal device. At this time, the terminal uses its built-in microphone to collect voice data and saves that voice data locally as a binary file.

[0344] Step 2:

[0345] The device sends the saved audio data to the server. The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. This process uses acoustic and language models to transform the audio signal into text. Text data is then generated.

[0346] Step 3:

[0347] The server passes the generated text data to the Google Cloud Natural Language API. The natural language processing engine analyzes the text and identifies the intent of the prompt. This process extracts keywords from the text and analyzes the context to understand the user's request. The intent is obtained as a result of the analysis.

[0348] Step 4:

[0349] The server consults a database to select appropriate instructions based on the user's intent. This database contains pre-prepared audio and visual guide templates. The server searches for and selects instruction content that matches the user's request. The instructions are then finalized as output.

[0350] Step 5:

[0351] The server transmits the selected instructions to the terminal as audio and visual media. The terminal plays these instructions and displays a visual demonstration on the screen. The user can perform the task according to the instructions by listening to the audio instructions and referring to the visual guide.

[0352] Step 6:

[0353] After completing an operation, the user enters feedback through an information terminal device. The terminal sends this feedback data to a server. The server analyzes the feedback and evaluates whether the problem has been resolved. This analysis involves determining the content of the feedback through text analysis and notifying local support personnel if additional support is needed.

[0354] This series of steps allows users to intuitively and efficiently manage and operate equipment using voice commands.

[0355] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0356] In the system of this invention, not only does the user instruct operations using voice commands, but the system also estimates the user's emotional state by analyzing the voice. This provides more personalized support and improves the user experience.

[0357] Users request support via devices such as smartphones or computers. After the device receives the user's voice command, it sends the voice data to the server. At this time, the server, equipped with an emotion engine, analyzes the voice data and processes it to identify the user's emotions. The emotion engine determines the user's emotions based on characteristics such as the intonation, speed, and pitch of the voice.

[0358] Based on the results of the emotion analysis, the server takes into account the user's emotional state and selects the most appropriate operation guide. For example, if the emotion engine determines that the user is confused, it aims to alleviate the user's anxiety by selecting a guide that provides detailed step-by-step instructions.

[0359] The user guide is tailored to the user's technical skill level, language, and emotional state, and is sent from the server to the terminal. This ensures the guide is tailored to the user's needs, enabling more effective support. Users can follow the instructions while viewing the audio and video displayed on the terminal.

[0360] Furthermore, users can provide feedback on the usefulness of the support, and this feedback is resent to the server via their device. The server analyzes the feedback and, if necessary, compares it with the results of the emotion engine to improve future support methods. Also, if the emotion engine determines from the feedback that the user's emotions remain unstable, it notifies local support staff to prioritize on-site assistance.

[0361] In this way, by utilizing an emotion engine, this system aims to not only solve technical problems but also comprehensively bridge the digital divide by taking into account the psychological aspects of the user.

[0362] The following describes the processing flow.

[0363] Step 1:

[0364] The user inputs a voice command into the device. The device records this using its microphone and temporarily stores it as audio data.

[0365] Step 2:

[0366] The device sends the recorded audio data to the server. This transmission is encrypted to prevent eavesdropping on the audio data.

[0367] Step 3:

[0368] The server receives the audio data and uses a speech recognition engine to convert the audio data into text.

[0369] Step 4:

[0370] The server analyzes the converted text using a natural language processing (NLP) engine to identify the type of support the user is requesting.

[0371] Step 5:

[0372] The server processes the voice data into an emotion engine, analyzing factors such as tone, intensity, and speed to estimate the user's emotional state.

[0373] Step 6:

[0374] The server combines the results of natural language processing and sentiment analysis to select the most appropriate operation guide from the database. This selection takes into account the user's pre-configured technical skill level and emotional state.

[0375] Step 7:

[0376] The server converts the selected operation guide into audio and video instructions and sends them to the terminal. These instructions are designed to be easy for the user to understand.

[0377] Step 8:

[0378] The device provides audio and video instructions to the user. The user can proceed with the operation while watching the on-screen demonstration.

[0379] Step 9:

[0380] Users can provide voice feedback. The device records that feedback and sends it to the server.

[0381] Step 10:

[0382] The server receives feedback and analyzes it using the emotion engine. If the user's emotions remain unstable, it notifies local support staff and prioritizes arranging on-site assistance.

[0383] (Example 2)

[0384] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0385] As modern information and communication technologies advance, users are increasingly required to operate devices based on diverse voice commands. However, conventional systems only perform simple voice recognition without considering the user's emotional state, making it difficult to provide effective support. Furthermore, they lack the flexibility to adapt to the diversity of users' technical abilities and languages. This can potentially expose users to stress and anxiety, potentially impairing their digital technology experience.

[0386] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0387] In this invention, the server includes a computing device means that performs speech recognition and emotion estimation to analyze voice information, a computing device means that selects and presents operational guidance based on the analysis results, and a computing device means that notifies local support personnel and arranges on-site support if the user's problem is not resolved. This enables the provision of personalized operational guides that correspond to the user's emotional state and flexible support that accommodates diversity in technical abilities and languages.

[0388] A "user" refers to an individual who operates the system and issues voice commands.

[0389] "Voice commands" refer to information that users give via voice commands or requests through their devices.

[0390] "Device means" refers to hardware or software functions for receiving voice commands from the user or displaying information.

[0391] "Voice information" refers to the voice data obtained after converting the user's voice commands into digital data.

[0392] "Speech recognition" refers to the technology that analyzes speech information and understands the content of speech commands.

[0393] "Emotion estimation" refers to a technology that determines a user's emotional state by analyzing voice information.

[0394] "Computation device means" refers to a computer system for speech recognition, emotion estimation, and the selection and presentation of operational guidance.

[0395] "Operational guidance" refers to information that provides guidelines and procedures for taking action based on the user's voice commands.

[0396] An "information display device" refers to hardware such as displays and projection equipment used to provide users with instructions on how to operate a system.

[0397] An "image device" refers to a device used to visually present information to a user.

[0398] A "regional support officer" refers to a specialist designated to provide on-site assistance in resolving users' technical problems.

[0399] "On-site support" refers to support activities that involve actually going to the user's location and directly resolving their problems.

[0400] "Multilingual support" refers to a feature that provides assistance in the language each user understands, for users with different language settings.

[0401] "Technical ability" refers to the technical knowledge and operational skills that a user possesses.

[0402] This invention is a system that improves the user experience by allowing users to operate it through voice commands. Users use a terminal such as a smartphone or computer to issue voice commands. The terminal converts the voice into digital data using a microphone or voice input module and transmits it to a server for analysis.

[0403] The server functions as a computing device equipped with advanced emotion estimation capabilities. This server analyzes audio data using speech recognition algorithms to estimate the user's emotional state. It analyzes features such as intonation, speed, and pitch, and utilizes generative AI models to determine the emotional state. For example, it builds emotion estimation models using machine learning frameworks such as TensorFlow and PyTorch.

[0404] Based on the analysis results, the server selects appropriate instructions. These instructions are customized based on the user's emotional state, technical ability, and multilingual support. The server-selected instructions are sent to the terminal, which then provides guidance to the user using a video display or audio output device.

[0405] As a concrete example, suppose a user gives a voice command saying, "I want to set up email on my new device." In this case, the server will assume the user is feeling anxious and select a detailed, step-by-step email setup guide for first-time users, sending it to the device. Alternatively, it can send a prompt to the generative AI model such as: "The user is anxious about setting up email on their new device. Please suggest a gentle and detailed email setup guide." This allows the user to proceed with confidence.

[0406] Furthermore, the server collects user feedback and evaluates the usefulness of voice guidance. Feedback analysis allows for continuous improvement of support services and responses to user requests. By notifying local support staff and arranging on-site assistance as needed, more comprehensive support becomes possible.

[0407] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0408] Step 1:

[0409] The user issues voice commands to the terminal. The input is the user's voice. The terminal uses its microphone to receive the voice as an analog signal. A voice processing module converts the analog signal into digital data, performing noise reduction and sound quality improvement. This digital data is output in the form of voice commands and is ready to be sent to the server.

[0410] Step 2:

[0411] The terminal sends digitized voice commands to the server via a security protocol. The input is digital voice data. The terminal performs data transfer, and the server receives the voice data. The output is the digital voice data transferred to the server.

[0412] Step 3:

[0413] The server analyzes the received digital audio data. The input is audio data. The server uses a speech recognition engine to convert it into text-based voice commands, and then utilizes an emotion estimation engine to determine the user's emotions. It analyzes data such as intonation, speed, and pitch, and outputs an emotion index.

[0414] Step 4:

[0415] The server selects the optimal operation guidance based on voice commands and emotion indicators. The inputs are voice commands and emotion indicators. Using a generative AI model, it infers the most appropriate operation guidance from the voice commands and further adjusts the guidance considering the emotion indicators. The output is customized operation guidance.

[0416] Step 5:

[0417] The server sends the configured operating instructions to the terminal. These instructions include video explanations and audio guides. The input is the customized operating instructions, which are sent to the terminal via data transfer from the server. The output is the operating instructions received by the terminal.

[0418] Step 6:

[0419] The terminal displays operating instructions received from the server to the user. The input is the operating instruction data. The terminal's display and speaker are used to present information to the user visually and audibly. The output is the operating instructions displayed to the user.

[0420] Step 7:

[0421] The user inputs feedback on the usefulness of the operation instructions into the terminal. This input is user feedback information. The terminal collects the input feedback and sends it to the server. The output of this transmission is the feedback data received by the server.

[0422] Step 8:

[0423] The server analyzes the received feedback and considers ways to improve future support. The input is feedback data. It compares this data with sentiment estimation results to evaluate system performance and identify areas for improvement. Furthermore, it sends notifications to local support staff as needed. The output is an improved support plan or support notification.

[0424] (Application Example 2)

[0425] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0426] Current content distribution services are insufficient in suggesting content that is tailored to the individual emotional state of users. Therefore, there is a need for services that are more attentive to users' emotions.

[0427] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0428] In this invention, the server includes means for receiving voice commands from the user and analyzing the voice data to estimate the user's emotional state; means for selecting an operation guide based on the results of voice recognition and emotion analysis and optimizing it according to the user's emotional state; and means for presenting the selected operation guide to the user visually and audibly. This enables personalized content suggestions based on the user's emotional state.

[0429] "Voice commands" refer to a method in which users give specific operational instructions to a system using their voice.

[0430] "Voice data" refers to the digital representation of voice input collected from users.

[0431] "Emotional state" refers to the emotional tendencies and psychological state analyzed from the user's voice.

[0432] "Speech recognition" is a technology that analyzes input speech data and converts it into text.

[0433] An "operation guide" is a set of instructions and procedures provided to a user to operate a system.

[0434] "Presenting information visually and aurally" refers to methods of conveying information to users using screen displays and audio output.

[0435] "Feedback" refers to the act of a user responding to or evaluating the services provided by a system.

[0436] "Emotional analysis" is a process that analyzes voice data to identify the user's emotional state.

[0437] "Improving support services" means optimizing the services and guides provided based on user feedback.

[0438] This invention is a personalized content delivery system that takes into account the user's emotional state, and primarily uses the user's voice input to analyze their emotions and provide optimal content. The system is implemented using the following means.

[0439] The system's terminal receives voice commands from the user. The voice data is sent to the server via the internet. The server converts the voice data into text using the Google Cloud Speech-to-Text API. Then, it uses Azure Text Analytics to perform sentiment analysis on this text data and estimate the user's emotional state.

[0440] Once the user's emotional state has been estimated, the server selects operation guides and content based on the results. The selection criteria include considering the user's psychological feedback and requests such as wanting to relax or become excited. During this process, generative AI models are used to optimize the user interface and operation experience.

[0441] Finally, the selected content and guides are presented to the device visually and aurally. Users can interact with these instructions and provide feedback. This feedback information is then sent back to the server to help improve the quality of future content delivery.

[0442] As a concrete example, consider a scenario where a user watches a movie and says, "Tell me what movie you recommend I watch next." If sentiment analysis of the voice data suggests the user is in a happy mood, a popular local comedy film will be suggested based on this. An example of a prompt would be, "Analyze the voice feedback the user has given on the content they watched and recommend content based on their emotional state."

[0443] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0444] Step 1:

[0445] The terminal receives voice commands from the user. Voice data is input and digitized as an audio file. The received voice data is transferred to a server via the intranet for processing.

[0446] Step 2:

[0447] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. This process transforms the audio data into a parseable text format, making it ready for use in subsequent processes.

[0448] Step 3:

[0449] The server uses Azure Text Analytics to perform sentiment analysis on the converted text data. Here, text data is input, and evaluation metrics indicating the user's emotional state are generated. This quantifies and outputs the user's psychological state.

[0450] Step 4:

[0451] The server selects appropriate content or user guides based on the sentiment analysis results. Here, sentiment evaluation metrics and the user's past history data are used as input, and a generative AI model is employed to select the optimal content. This output forms the basis for the next step.

[0452] Step 5:

[0453] The selected content or operating guide is sent to the device. The content is entered into the device and presented to the user visually and aurally. Through this process, the user receives personalized content.

[0454] Step 6:

[0455] The user provides voice feedback on the content presented by the system. The device then records the user's feedback again as voice data and prepares to send it to the server.

[0456] Step 7:

[0457] The server receives feedback from the user and performs sentiment analysis again. Using the feedback audio data as input, it outputs a new sentiment evaluation metric. This allows the server to confirm user satisfaction and use the results to improve service quality.

[0458] Step 8:

[0459] Based on feedback, the server improves its content selection algorithm and sentiment analysis model as needed. This improvement process uses accumulated sentiment data and feedback as input, enabling the delivery of more accurate content.

[0460] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0461] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0462] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0463] [Third Embodiment]

[0464] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0465] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0466] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0467] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0468] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0469] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0470] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0471] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0472] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0473] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0474] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0475] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0476] The system of this invention supports bridging the digital divide by allowing users to give voice commands for operation. Users use terminals such as smartphones or computers to inquire by voice about operations that require support.

[0477] After receiving a voice command from the user, the terminal sends the data to the server via the internet. The server converts the received voice data into text using a speech recognition engine. Next, it uses a natural language processing (NLP) engine to analyze the user's request and identify what kind of support is needed.

[0478] Based on the analysis results, the server selects the appropriate guide from its stored operation guides. These guides include audio and video instructions and are presented in a highly user-friendly format. The selected guide is tailored to the user's skill level and language settings.

[0479] Audio and video content generated by the server is sent to the terminal and provided to the user. The terminal supports the user in easily understanding and performing operations by displaying visual demonstrations on the screen along with voice instructions. For example, if the user asks by voice, "Tell me how to send an email," the server provides a series of video guides from basic settings in the email application to creating and sending an email.

[0480] When a user provides feedback, the device sends that information back to the server. The server analyzes the feedback, evaluates whether the user's problem has been resolved, and provides additional support as needed. If the user's problem is technically complex and cannot be resolved, the server can notify local support staff and schedule an on-site assistance visit.

[0481] This system features multilingual support and customization options tailored to different skill levels, enabling it to provide equally accurate support to users with diverse backgrounds. Its aim is to assist individuals who feel uneasy about using digital devices and effectively bridge the digital divide.

[0482] The following describes the processing flow.

[0483] Step 1:

[0484] The user issues voice commands to a device such as a smartphone or PC. The commands are recorded by the microphone built into the device.

[0485] Step 2:

[0486] The device sends the recorded audio data to the server via the internet. This communication is encrypted and conducted in a secure manner.

[0487] Step 3:

[0488] The server processes the received audio data through a speech recognition engine, converting the speech into text. Speech recognition is performed with high accuracy using an AI model.

[0489] Step 4:

[0490] The server analyzes the converted text using a natural language processing (NLP) engine to identify the type of support the user is seeking. This analysis helps understand the user's intent and the context of their question.

[0491] Step 5:

[0492] Based on the analysis results, the server selects the appropriate operation guide from the database. This guide is customized according to the user's technical skill level and language settings.

[0493] Step 6:

[0494] The server generates audio and video instructions based on the selected operation guide and sends them to the terminal. These instructions are presented in a format that is easy for the user to understand.

[0495] Step 7:

[0496] The device displays audio and video instructions received by the user. In addition to audio explanations, visual demonstrations are provided on the screen.

[0497] Step 8:

[0498] The system can perform actions based on the information provided by the user and provide feedback on whether the support was helpful. This feedback may also be provided via voice.

[0499] Step 9:

[0500] The device sends user feedback back to the server. Based on this feedback, the server determines whether to improve support or if additional support is needed.

[0501] Step 10:

[0502] If the user's problem is not resolved, the server will notify local support staff. On-site support will be arranged depending on the situation.

[0503] (Example 1)

[0504] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0505] With the evolution of digital technology, a problem arises where users with limited technical skills and knowledge are unable to effectively utilize digital devices and services. This digital divide is particularly pronounced in multilingual environments and among individuals with diverse technical skill sets, hindering access to information. Furthermore, there is the challenge of difficulty in receiving prompt and appropriate support when technical assistance is needed.

[0506] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0507] In this invention, the server includes a computing device means for performing automatic speech recognition to convert speech data into text information, a language analysis means for analyzing the text information and identifying the requested support, and a computing device means for selecting and providing appropriate operating guidelines based on the analysis results. This enables users to easily operate digital devices and services and receive necessary support in multiple languages ​​and adaptively, regardless of their technical skills or knowledge.

[0508] "Computer means" refers to a device that has the function of receiving voice instructions from the user and presenting selected operating guidelines in both audio and video.

[0509] The "computational device means" is a device that performs automatic speech recognition to convert speech data into text information, and further analyzes the text information to identify the requested support content.

[0510] "Language analysis means" refers to technology that analyzes text data converted by speech recognition to understand user requests.

[0511] "Operating guidelines" are audio and video content that explains the procedures for specific operations or tasks, providing effective guidance to users.

[0512] A "technical support provider" is a specialist or staff member who provides assistance when technical problems cannot be resolved.

[0513] "On-site support" refers to technical support personnel traveling to the user's location to provide direct assistance.

[0514] "Multilingual processing" is a technology that has the capability to provide appropriate support in each language to users who use different languages.

[0515] "Adaptation" refers to the process by which a system adjusts itself to provide optimal support based on the user's technical skill level.

[0516] The system of the present invention aims to bridge the digital divide by enabling users to request the operation of digital devices using voice commands. Users input voice commands using terminal devices such as smartphones or computers, and this voice data is transmitted to a server.

[0517] The server uses a speech recognition engine to process the received audio data. This engine can be a commonly used speech recognition software, such as the Google Speech-to-Text API. This converts the audio data into text information.

[0518] Next, the server uses a natural language processing (NLP) engine to analyze the transcribed data. This allows it to accurately understand the intent of the user's request and identify the necessary assistance. Common natural language processing tools can be used for this analysis.

[0519] Once the analysis is complete, the server selects the appropriate guide from several operation guides stored in storage and configures its contents. This guide includes detailed audio and visual instructions, making it easy for the user to understand and perform the operation. Customization is also possible according to the user's technical skill level and language settings, and the user interface is adjusted based on the language and skills selected by the user.

[0520] The device provides the user with audio and video guides received from the server. This device displays visual demonstrations and, in conjunction with audio instructions, assists the user in smooth operation. For example, if the user asks "How to create a homepage," the server provides a beginner-friendly HTML / CSS basics guide and plays a visualized version of that guide on the device.

[0521] Furthermore, when a user submits feedback, this feedback information is sent back to the server, which analyzes it and evaluates whether the user's problem has been resolved. If the problem is difficult to resolve, the server notifies local technical support staff, and on-site assistance is arranged.

[0522] Example prompt: "Create a beginner's guide to HTML / CSS basics based on the user's voice command 'How to create a homepage'."

[0523] This allows users to efficiently utilize digital devices regardless of their technical knowledge.

[0524] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0525] Step 1:

[0526] The user inputs voice commands using the device. Specifically, they use the device's microphone to issue voice commands such as "I want to edit the photo." The input is the user's voice command, and the output is audio data.

[0527] Step 2:

[0528] The terminal sends the acquired audio data directly to the server. The input is the user's audio data, and the output is the audio data that the server uses to begin processing. The transmission process involves data transfer from the terminal to the server.

[0529] Step 3:

[0530] The server converts the received audio data into text information using a speech recognition engine. In this process, the input is audio data and the output is text data. The speech recognition engine analyzes the audio data and generates the corresponding text.

[0531] Step 4:

[0532] The server uses the converted text data to perform natural language processing (NLP) and analyze the user's instructions. The input is text data, and the output is the analysis result. The NLP engine understands the intent of the text and identifies the necessary support.

[0533] Step 5:

[0534] The server selects an operation guide based on the analysis results. The input is the analysis results, and the output is the selected operation guide. The server selects the most appropriate guide from its internal database.

[0535] Step 6:

[0536] The server customizes the selected operation guide to match the user's skill level and language. The input is the operation guide and user configuration information, and the output is the customized guide. The server adjusts the guide content to suit the user.

[0537] Step 7:

[0538] The server transmits customized audio and video content to the terminal. The input is the adjusted guide content, and the output is the content available to the user. The content is delivered to the user's terminal via data communication.

[0539] Step 8:

[0540] The terminal provides the user with received audio and video guides and plays visual demonstrations on the screen. The input is guide content from the server, and the output is displayed content that facilitates user understanding. The terminal provides an environment where it can view the content, supporting the user experience.

[0541] Step 9:

[0542] The user views the guide provided on the terminal, performs the operations, and sends feedback to the terminal as needed. The input is the user's operational feedback, and the output is feedback data. This information is then sent back to the server for further analysis.

[0543] (Application Example 1)

[0544] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0545] There is a need to overcome the difficulties faced by individuals unfamiliar with digital technology when efficiently using safety devices and information processing systems in homes and workplaces, and to provide support tailored to diverse language and technical skill levels. The goal is to enable accurate and rapid responses in emergencies and to allow for the safe use of equipment.

[0546] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0547] In this invention, the server includes means including an information terminal device for receiving voice commands from a user; computer means for performing voice recognition processing for analyzing voice data; computer means for selecting and providing instructions based on the analysis results; means including an information terminal device for presenting the selected instructions to the user audibly and visually; means including an information terminal device for receiving user feedback and transmitting it to the computer; computer means for notifying local support staff and arranging on-site support if the user's problem is not resolved; means for managing safety protection devices based on voice operation; and means for providing voice guidance in emergencies. This enables even users unfamiliar with digital technology to intuitively and quickly operate safety protection devices and other information systems.

[0548] An "information terminal device" is a device that receives voice commands from a user and processes them in an appropriate manner.

[0549] "Computing means" refers to a computing device that has functions for processing digital data, such as speech recognition and analysis of instructions.

[0550] "Speech recognition processing" is the process of analyzing audio data received from a user and converting it into text data.

[0551] A "safety protection device" is a device that has security functions to protect people or assets.

[0552] "Means of providing voice guidance in emergencies" refers to a function that provides voice guidance to encourage users to take appropriate action when an emergency occurs.

[0553] A "local support provider" refers to a support provider who can visit users in person to provide assistance in order to solve their technical problems.

[0554] "User feedback" refers to the process of users providing results and opinions on their experiences, and includes the system receiving and analyzing this feedback.

[0555] "Instructions" refers to operation guides and audio / video instructions provided based on user requests or questions.

[0556] The embodiments for carrying out the invention are described below.

[0557] This invention provides a system that allows users to manage and operate devices by inputting voice commands through an information terminal device and using voice recognition. First, the user sends a voice command from the information terminal device. The terminal collects this command as voice data and transmits it to a computing means. The computing means, in particular, voice recognition software such as the Google Cloud Speech-to-Text API, converts this into text data and further analyzes the content of the instruction using the Google Cloud Natural Language API.

[0558] Based on the analysis results, appropriate instructions are sent back to the information terminal device. The user receives the instructions via voice and visual interfaces and, based on them, can, for example, remotely control the settings of a home safety device by voice. User feedback is also sent back from the terminal to the computer, and arrangements are made to notify local support staff if the user's problem is not resolved.

[0559] This system includes a concrete example where the settings of a security camera are automatically adjusted when the user inputs a prompt as a voice command, such as "I want to be notified when my child comes home." The aim is to bridge the digital divide when managing and operating safety devices through such voice commands.

[0560] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0561] Step 1:

[0562] The user inputs voice commands into the information terminal device. At this time, the terminal uses its built-in microphone to collect voice data and saves that voice data locally as a binary file.

[0563] Step 2:

[0564] The device sends the saved audio data to the server. The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. This process uses acoustic and language models to transform the audio signal into text. Text data is then generated.

[0565] Step 3:

[0566] The server passes the generated text data to the Google Cloud Natural Language API. The natural language processing engine analyzes the text and identifies the intent of the prompt. This process extracts keywords from the text and analyzes the context to understand the user's request. The intent is obtained as a result of the analysis.

[0567] Step 4:

[0568] The server consults a database to select appropriate instructions based on the user's intent. This database contains pre-prepared audio and visual guide templates. The server searches for and selects instruction content that matches the user's request. The instructions are then finalized as output.

[0569] Step 5:

[0570] The server transmits the selected instructions to the terminal as audio and visual media. The terminal plays these instructions and displays a visual demonstration on the screen. The user can perform the task according to the instructions by listening to the audio instructions and referring to the visual guide.

[0571] Step 6:

[0572] After completing an operation, the user enters feedback through an information terminal device. The terminal sends this feedback data to a server. The server analyzes the feedback and evaluates whether the problem has been resolved. This analysis involves determining the content of the feedback through text analysis and notifying local support personnel if additional support is needed.

[0573] This series of steps allows users to intuitively and efficiently manage and operate equipment using voice commands.

[0574] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0575] In the system of this invention, not only does the user instruct operations using voice commands, but the system also estimates the user's emotional state by analyzing the voice. This provides more personalized support and improves the user experience.

[0576] Users request support via devices such as smartphones or computers. After the device receives the user's voice command, it sends the voice data to the server. At this time, the server, equipped with an emotion engine, analyzes the voice data and processes it to identify the user's emotions. The emotion engine determines the user's emotions based on characteristics such as the intonation, speed, and pitch of the voice.

[0577] Based on the results of the emotion analysis, the server takes into account the user's emotional state and selects the most appropriate operation guide. For example, if the emotion engine determines that the user is confused, it aims to alleviate the user's anxiety by selecting a guide that provides detailed step-by-step instructions.

[0578] The user guide is tailored to the user's technical skill level, language, and emotional state, and is sent from the server to the terminal. This ensures the guide is tailored to the user's needs, enabling more effective support. Users can follow the instructions while viewing the audio and video displayed on the terminal.

[0579] Furthermore, users can provide feedback on the usefulness of the support, and this feedback is resent to the server via their device. The server analyzes the feedback and, if necessary, compares it with the results of the emotion engine to improve future support methods. Also, if the emotion engine determines from the feedback that the user's emotions remain unstable, it notifies local support staff to prioritize on-site assistance.

[0580] In this way, by utilizing an emotion engine, this system aims to not only solve technical problems but also comprehensively bridge the digital divide by taking into account the psychological aspects of the user.

[0581] The following describes the processing flow.

[0582] Step 1:

[0583] The user inputs a voice command into the device. The device records this using its microphone and temporarily stores it as audio data.

[0584] Step 2:

[0585] The device sends the recorded audio data to the server. This transmission is encrypted to prevent eavesdropping on the audio data.

[0586] Step 3:

[0587] The server receives the audio data and uses a speech recognition engine to convert the audio data into text.

[0588] Step 4:

[0589] The server analyzes the converted text using a natural language processing (NLP) engine to identify the type of support the user is requesting.

[0590] Step 5:

[0591] The server processes the voice data into an emotion engine, analyzing factors such as tone, intensity, and speed to estimate the user's emotional state.

[0592] Step 6:

[0593] The server combines the results of natural language processing and sentiment analysis to select the most appropriate operation guide from the database. This selection takes into account the user's pre-configured technical skill level and emotional state.

[0594] Step 7:

[0595] The server converts the selected operation guide into audio and video instructions and sends them to the terminal. These instructions are designed to be easy for the user to understand.

[0596] Step 8:

[0597] The device provides audio and video instructions to the user. The user can proceed with the operation while watching the on-screen demonstration.

[0598] Step 9:

[0599] Users can provide voice feedback. The device records that feedback and sends it to the server.

[0600] Step 10:

[0601] The server receives feedback and analyzes it using the emotion engine. If the user's emotions remain unstable, it notifies local support staff and prioritizes arranging on-site assistance.

[0602] (Example 2)

[0603] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0604] As modern information and communication technologies advance, users are increasingly required to operate devices based on diverse voice commands. However, conventional systems only perform simple voice recognition without considering the user's emotional state, making it difficult to provide effective support. Furthermore, they lack the flexibility to adapt to the diversity of users' technical abilities and languages. This can potentially expose users to stress and anxiety, potentially impairing their digital technology experience.

[0605] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0606] In this invention, the server includes a computing device means that performs speech recognition and emotion estimation to analyze voice information, a computing device means that selects and presents operational guidance based on the analysis results, and a computing device means that notifies local support personnel and arranges on-site support if the user's problem is not resolved. This enables the provision of personalized operational guides that correspond to the user's emotional state and flexible support that accommodates diversity in technical abilities and languages.

[0607] A "user" refers to an individual who operates the system and issues voice commands.

[0608] "Voice commands" refer to information that users give via voice commands or requests through their devices.

[0609] "Device means" refers to hardware or software functions for receiving voice commands from the user or displaying information.

[0610] "Voice information" refers to the voice data obtained after converting the user's voice commands into digital data.

[0611] "Speech recognition" refers to the technology that analyzes speech information and understands the content of speech commands.

[0612] "Emotion estimation" refers to a technology that determines a user's emotional state by analyzing voice information.

[0613] "Computation device means" refers to a computer system for speech recognition, emotion estimation, and the selection and presentation of operational guidance.

[0614] "Operational guidance" refers to information that provides guidelines and procedures for taking action based on the user's voice commands.

[0615] An "information display device" refers to hardware such as displays and projection equipment used to provide users with instructions on how to operate a system.

[0616] An "image device" refers to a device used to visually present information to a user.

[0617] A "regional support officer" refers to a specialist designated to provide on-site assistance in resolving users' technical problems.

[0618] "On-site support" refers to support activities that involve actually going to the user's location and directly resolving their problems.

[0619] "Multilingual support" refers to a feature that provides assistance in the language each user understands, for users with different language settings.

[0620] "Technical ability" refers to the technical knowledge and operational skills that a user possesses.

[0621] This invention is a system that improves the user experience by allowing users to operate it through voice commands. Users use a terminal such as a smartphone or computer to issue voice commands. The terminal converts the voice into digital data using a microphone or voice input module and transmits it to a server for analysis.

[0622] The server functions as a computing device equipped with advanced emotion estimation capabilities. This server analyzes audio data using speech recognition algorithms to estimate the user's emotional state. It analyzes features such as intonation, speed, and pitch, and utilizes generative AI models to determine the emotional state. For example, it builds emotion estimation models using machine learning frameworks such as TensorFlow and PyTorch.

[0623] Based on the analysis results, the server selects appropriate instructions. These instructions are customized based on the user's emotional state, technical ability, and multilingual support. The server-selected instructions are sent to the terminal, which then provides guidance to the user using a video display or audio output device.

[0624] As a concrete example, suppose a user gives a voice command saying, "I want to set up email on my new device." In this case, the server will assume the user is feeling anxious and select a detailed, step-by-step email setup guide for first-time users, sending it to the device. Alternatively, it can send a prompt to the generative AI model such as: "The user is anxious about setting up email on their new device. Please suggest a gentle and detailed email setup guide." This allows the user to proceed with confidence.

[0625] Furthermore, the server collects user feedback and evaluates the usefulness of voice guidance. Feedback analysis allows for continuous improvement of support services and responses to user requests. By notifying local support staff and arranging on-site assistance as needed, more comprehensive support becomes possible.

[0626] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0627] Step 1:

[0628] The user issues voice commands to the terminal. The input is the user's voice. The terminal uses its microphone to receive the voice as an analog signal. A voice processing module converts the analog signal into digital data, performing noise reduction and sound quality improvement. This digital data is output in the form of voice commands and is ready to be sent to the server.

[0629] Step 2:

[0630] The terminal sends digitized voice commands to the server via a security protocol. The input is digital voice data. The terminal performs data transfer, and the server receives the voice data. The output is the digital voice data transferred to the server.

[0631] Step 3:

[0632] The server analyzes the received digital audio data. The input is audio data. The server uses a speech recognition engine to convert it into text-based voice commands, and then utilizes an emotion estimation engine to determine the user's emotions. It analyzes data such as intonation, speed, and pitch, and outputs an emotion index.

[0633] Step 4:

[0634] The server selects the optimal operation guidance based on voice commands and emotion indicators. The inputs are voice commands and emotion indicators. Using a generative AI model, it infers the most appropriate operation guidance from the voice commands and further adjusts the guidance considering the emotion indicators. The output is customized operation guidance.

[0635] Step 5:

[0636] The server sends the configured operating instructions to the terminal. These instructions include video explanations and audio guides. The input is the customized operating instructions, which are sent to the terminal via data transfer from the server. The output is the operating instructions received by the terminal.

[0637] Step 6:

[0638] The terminal displays operating instructions received from the server to the user. The input is the operating instruction data. The terminal's display and speaker are used to present information to the user visually and audibly. The output is the operating instructions displayed to the user.

[0639] Step 7:

[0640] The user inputs feedback on the usefulness of the operation instructions into the terminal. This input is user feedback information. The terminal collects the input feedback and sends it to the server. The output of this transmission is the feedback data received by the server.

[0641] Step 8:

[0642] The server analyzes the received feedback and considers ways to improve future support. The input is feedback data. It compares this data with sentiment estimation results to evaluate system performance and identify areas for improvement. Furthermore, it sends notifications to local support staff as needed. The output is an improved support plan or support notification.

[0643] (Application Example 2)

[0644] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0645] Current content distribution services are insufficient in suggesting content that is tailored to the individual emotional state of users. Therefore, there is a need for services that are more attentive to users' emotions.

[0646] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0647] In this invention, the server includes means for receiving voice commands from the user and analyzing the voice data to estimate the user's emotional state; means for selecting an operation guide based on the results of voice recognition and emotion analysis and optimizing it according to the user's emotional state; and means for presenting the selected operation guide to the user visually and audibly. This enables personalized content suggestions based on the user's emotional state.

[0648] "Voice commands" refer to a method in which users give specific operational instructions to a system using their voice.

[0649] "Voice data" refers to the digital representation of voice input collected from users.

[0650] "Emotional state" refers to the emotional tendencies and psychological state analyzed from the user's voice.

[0651] "Speech recognition" is a technology that analyzes input speech data and converts it into text.

[0652] An "operation guide" is a set of instructions and procedures provided to a user to operate a system.

[0653] "Presenting information visually and aurally" refers to methods of conveying information to users using screen displays and audio output.

[0654] "Feedback" refers to the act of a user responding to or evaluating the services provided by a system.

[0655] "Emotional analysis" is a process that analyzes voice data to identify the user's emotional state.

[0656] "Improving support services" means optimizing the services and guides provided based on user feedback.

[0657] This invention is a personalized content delivery system that takes into account the user's emotional state, and primarily uses the user's voice input to analyze their emotions and provide optimal content. The system is implemented using the following means.

[0658] The system's terminal receives voice commands from the user. The voice data is sent to the server via the internet. The server converts the voice data into text using the Google Cloud Speech-to-Text API. Then, it uses Azure Text Analytics to perform sentiment analysis on this text data and estimate the user's emotional state.

[0659] Once the user's emotional state has been estimated, the server selects operation guides and content based on the results. The selection criteria include considering the user's psychological feedback and requests such as wanting to relax or become excited. During this process, generative AI models are used to optimize the user interface and operation experience.

[0660] Finally, the selected content and guides are presented to the device visually and aurally. Users can interact with these instructions and provide feedback. This feedback information is then sent back to the server to help improve the quality of future content delivery.

[0661] As a concrete example, consider a scenario where a user watches a movie and says, "Tell me what movie you recommend I watch next." If sentiment analysis of the voice data suggests the user is in a happy mood, a popular local comedy film will be suggested based on this. An example of a prompt would be, "Analyze the voice feedback the user has given on the content they watched and recommend content based on their emotional state."

[0662] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0663] Step 1:

[0664] The terminal receives voice commands from the user. Voice data is input and digitized as an audio file. The received voice data is transferred to a server via the intranet for processing.

[0665] Step 2:

[0666] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. This process transforms the audio data into a parseable text format, making it ready for use in subsequent processes.

[0667] Step 3:

[0668] The server uses Azure Text Analytics to perform sentiment analysis on the converted text data. Here, text data is input, and evaluation metrics indicating the user's emotional state are generated. This quantifies and outputs the user's psychological state.

[0669] Step 4:

[0670] The server selects appropriate content or user guides based on the sentiment analysis results. Here, sentiment evaluation metrics and the user's past history data are used as input, and a generative AI model is employed to select the optimal content. This output forms the basis for the next step.

[0671] Step 5:

[0672] The selected content or operating guide is sent to the device. The content is entered into the device and presented to the user visually and aurally. Through this process, the user receives personalized content.

[0673] Step 6:

[0674] The user provides voice feedback on the content presented by the system. The device then records the user's feedback again as voice data and prepares to send it to the server.

[0675] Step 7:

[0676] The server receives feedback from the user and performs sentiment analysis again. Using the feedback audio data as input, it outputs a new sentiment evaluation metric. This allows the server to confirm user satisfaction and use the results to improve service quality.

[0677] Step 8:

[0678] Based on feedback, the server improves its content selection algorithm and sentiment analysis model as needed. This improvement process uses accumulated sentiment data and feedback as input, enabling the delivery of more accurate content.

[0679] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0680] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0681] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0682] [Fourth Embodiment]

[0683] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0684] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0685] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0686] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0687] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0688] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0689] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0690] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0691] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0692] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0693] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0694] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0695] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0696] The system of this invention supports bridging the digital divide by allowing users to give voice commands for operation. Users use terminals such as smartphones or computers to inquire by voice about operations that require support.

[0697] After receiving a voice command from the user, the terminal sends the data to the server via the internet. The server converts the received voice data into text using a speech recognition engine. Next, it uses a natural language processing (NLP) engine to analyze the user's request and identify what kind of support is needed.

[0698] Based on the analysis results, the server selects the appropriate guide from its stored operation guides. These guides include audio and video instructions and are presented in a highly user-friendly format. The selected guide is tailored to the user's skill level and language settings.

[0699] Audio and video content generated by the server is sent to the terminal and provided to the user. The terminal supports the user in easily understanding and performing operations by displaying visual demonstrations on the screen along with voice instructions. For example, if the user asks by voice, "Tell me how to send an email," the server provides a series of video guides from basic settings in the email application to creating and sending an email.

[0700] When a user provides feedback, the device sends that information back to the server. The server analyzes the feedback, evaluates whether the user's problem has been resolved, and provides additional support as needed. If the user's problem is technically complex and cannot be resolved, the server can notify local support staff and schedule an on-site assistance visit.

[0701] This system features multilingual support and customization options tailored to different skill levels, enabling it to provide equally accurate support to users with diverse backgrounds. Its aim is to assist individuals who feel uneasy about using digital devices and effectively bridge the digital divide.

[0702] The following describes the processing flow.

[0703] Step 1:

[0704] The user issues voice commands to a device such as a smartphone or PC. The commands are recorded by the microphone built into the device.

[0705] Step 2:

[0706] The device sends the recorded audio data to the server via the internet. This communication is encrypted and conducted in a secure manner.

[0707] Step 3:

[0708] The server processes the received audio data through a speech recognition engine, converting the speech into text. Speech recognition is performed with high accuracy using an AI model.

[0709] Step 4:

[0710] The server analyzes the converted text using a natural language processing (NLP) engine to identify the type of support the user is seeking. This analysis helps understand the user's intent and the context of their question.

[0711] Step 5:

[0712] Based on the analysis results, the server selects the appropriate operation guide from the database. This guide is customized according to the user's technical skill level and language settings.

[0713] Step 6:

[0714] The server generates audio and video instructions based on the selected operation guide and sends them to the terminal. These instructions are presented in a format that is easy for the user to understand.

[0715] Step 7:

[0716] The device displays audio and video instructions received by the user. In addition to audio explanations, visual demonstrations are provided on the screen.

[0717] Step 8:

[0718] The system can perform actions based on the information provided by the user and provide feedback on whether the support was helpful. This feedback may also be provided via voice.

[0719] Step 9:

[0720] The device sends user feedback back to the server. Based on this feedback, the server determines whether to improve support or if additional support is needed.

[0721] Step 10:

[0722] If the user's problem is not resolved, the server will notify local support staff. On-site support will be arranged depending on the situation.

[0723] (Example 1)

[0724] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0725] With the evolution of digital technology, a problem arises where users with limited technical skills and knowledge are unable to effectively utilize digital devices and services. This digital divide is particularly pronounced in multilingual environments and among individuals with diverse technical skill sets, hindering access to information. Furthermore, there is the challenge of difficulty in receiving prompt and appropriate support when technical assistance is needed.

[0726] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0727] In this invention, the server includes a computing device means for performing automatic speech recognition to convert speech data into text information, a language analysis means for analyzing the text information and identifying the requested support, and a computing device means for selecting and providing appropriate operating guidelines based on the analysis results. This enables users to easily operate digital devices and services and receive necessary support in multiple languages ​​and adaptively, regardless of their technical skills or knowledge.

[0728] "Computer means" refers to a device that has the function of receiving voice instructions from the user and presenting selected operating guidelines in both audio and video.

[0729] The "computational device means" is a device that performs automatic speech recognition to convert speech data into text information, and further analyzes the text information to identify the requested support content.

[0730] "Language analysis means" refers to technology that analyzes text data converted by speech recognition to understand user requests.

[0731] "Operating guidelines" are audio and video content that explains the procedures for specific operations or tasks, providing effective guidance to users.

[0732] A "technical support provider" is a specialist or staff member who provides assistance when technical problems cannot be resolved.

[0733] "On-site support" refers to technical support personnel traveling to the user's location to provide direct assistance.

[0734] "Multilingual processing" is a technology that has the capability to provide appropriate support in each language to users who use different languages.

[0735] "Adaptation" refers to the process by which a system adjusts itself to provide optimal support based on the user's technical skill level.

[0736] The system of the present invention aims to bridge the digital divide by enabling users to request the operation of digital devices using voice commands. Users input voice commands using terminal devices such as smartphones or computers, and this voice data is transmitted to a server.

[0737] The server uses a speech recognition engine to process the received audio data. This engine can be a commonly used speech recognition software, such as the Google Speech-to-Text API. This converts the audio data into text information.

[0738] Next, the server uses a natural language processing (NLP) engine to analyze the transcribed data. This allows it to accurately understand the intent of the user's request and identify the necessary assistance. Common natural language processing tools can be used for this analysis.

[0739] Once the analysis is complete, the server selects the appropriate guide from several operation guides stored in storage and configures its contents. This guide includes detailed audio and visual instructions, making it easy for the user to understand and perform the operation. Customization is also possible according to the user's technical skill level and language settings, and the user interface is adjusted based on the language and skills selected by the user.

[0740] The device provides the user with audio and video guides received from the server. This device displays visual demonstrations and, in conjunction with audio instructions, assists the user in smooth operation. For example, if the user asks "How to create a homepage," the server provides a beginner-friendly HTML / CSS basics guide and plays a visualized version of that guide on the device.

[0741] Furthermore, when a user submits feedback, this feedback information is sent back to the server, which analyzes it and evaluates whether the user's problem has been resolved. If the problem is difficult to resolve, the server notifies local technical support staff, and on-site assistance is arranged.

[0742] Example prompt: "Create a beginner's guide to HTML / CSS basics based on the user's voice command 'How to create a homepage'."

[0743] This allows users to efficiently utilize digital devices regardless of their technical knowledge.

[0744] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0745] Step 1:

[0746] The user inputs voice commands using the device. Specifically, they use the device's microphone to issue voice commands such as "I want to edit the photo." The input is the user's voice command, and the output is audio data.

[0747] Step 2:

[0748] The terminal sends the acquired audio data directly to the server. The input is the user's audio data, and the output is the audio data that the server uses to begin processing. The transmission process involves data transfer from the terminal to the server.

[0749] Step 3:

[0750] The server converts the received audio data into text information using a speech recognition engine. In this process, the input is audio data and the output is text data. The speech recognition engine analyzes the audio data and generates the corresponding text.

[0751] Step 4:

[0752] The server uses the converted text data to perform natural language processing (NLP) and analyze the user's instructions. The input is text data, and the output is the analysis result. The NLP engine understands the intent of the text and identifies the necessary support.

[0753] Step 5:

[0754] The server selects an operation guide based on the analysis results. The input is the analysis results, and the output is the selected operation guide. The server selects the most appropriate guide from its internal database.

[0755] Step 6:

[0756] The server customizes the selected operation guide to match the user's skill level and language. The input is the operation guide and user configuration information, and the output is the customized guide. The server adjusts the guide content to suit the user.

[0757] Step 7:

[0758] The server transmits customized audio and video content to the terminal. The input is the adjusted guide content, and the output is the content available to the user. The content is delivered to the user's terminal via data communication.

[0759] Step 8:

[0760] The terminal provides the user with received audio and video guides and plays visual demonstrations on the screen. The input is guide content from the server, and the output is displayed content that facilitates user understanding. The terminal provides an environment where it can view the content, supporting the user experience.

[0761] Step 9:

[0762] The user views the guide provided on the terminal, performs the operations, and sends feedback to the terminal as needed. The input is the user's operational feedback, and the output is feedback data. This information is then sent back to the server for further analysis.

[0763] (Application Example 1)

[0764] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0765] There is a need to overcome the difficulties faced by individuals unfamiliar with digital technology when efficiently using safety devices and information processing systems in homes and workplaces, and to provide support tailored to diverse language and technical skill levels. The goal is to enable accurate and rapid responses in emergencies and to allow for the safe use of equipment.

[0766] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0767] In this invention, the server includes means including an information terminal device for receiving voice commands from a user; computer means for performing voice recognition processing for analyzing voice data; computer means for selecting and providing instructions based on the analysis results; means including an information terminal device for presenting the selected instructions to the user audibly and visually; means including an information terminal device for receiving user feedback and transmitting it to the computer; computer means for notifying local support staff and arranging on-site support if the user's problem is not resolved; means for managing safety protection devices based on voice operation; and means for providing voice guidance in emergencies. This enables even users unfamiliar with digital technology to intuitively and quickly operate safety protection devices and other information systems.

[0768] An "information terminal device" is a device that receives voice commands from a user and processes them in an appropriate manner.

[0769] "Computing means" refers to a computing device that has functions for processing digital data, such as speech recognition and analysis of instructions.

[0770] "Speech recognition processing" is the process of analyzing audio data received from a user and converting it into text data.

[0771] A "safety protection device" is a device that has security functions to protect people or assets.

[0772] "Means of providing voice guidance in emergencies" refers to a function that provides voice guidance to encourage users to take appropriate action when an emergency occurs.

[0773] A "local support provider" refers to a support provider who can visit users in person to provide assistance in order to solve their technical problems.

[0774] "User feedback" refers to the process of users providing results and opinions on their experiences, and includes the system receiving and analyzing this feedback.

[0775] "Instructions" refers to operation guides and audio / video instructions provided based on user requests or questions.

[0776] The embodiments for carrying out the invention are described below.

[0777] This invention provides a system that allows users to manage and operate devices by inputting voice commands through an information terminal device and using voice recognition. First, the user sends a voice command from the information terminal device. The terminal collects this command as voice data and transmits it to a computing means. The computing means, in particular, voice recognition software such as the Google Cloud Speech-to-Text API, converts this into text data and further analyzes the content of the instruction using the Google Cloud Natural Language API.

[0778] Based on the analysis results, appropriate instructions are sent back to the information terminal device. The user receives the instructions via voice and visual interfaces and, based on them, can, for example, remotely control the settings of a home safety device by voice. User feedback is also sent back from the terminal to the computer, and arrangements are made to notify local support staff if the user's problem is not resolved.

[0779] This system includes a concrete example where the settings of a security camera are automatically adjusted when the user inputs a prompt as a voice command, such as "I want to be notified when my child comes home." The aim is to bridge the digital divide when managing and operating safety devices through such voice commands.

[0780] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0781] Step 1:

[0782] The user inputs voice commands into the information terminal device. At this time, the terminal uses its built-in microphone to collect voice data and saves that voice data locally as a binary file.

[0783] Step 2:

[0784] The device sends the saved audio data to the server. The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. This process uses acoustic and language models to transform the audio signal into text. Text data is then generated.

[0785] Step 3:

[0786] The server passes the generated text data to the Google Cloud Natural Language API. The natural language processing engine analyzes the text and identifies the intent of the prompt. This process extracts keywords from the text and analyzes the context to understand the user's request. The intent is obtained as a result of the analysis.

[0787] Step 4:

[0788] The server consults a database to select appropriate instructions based on the user's intent. This database contains pre-prepared audio and visual guide templates. The server searches for and selects instruction content that matches the user's request. The instructions are then finalized as output.

[0789] Step 5:

[0790] The server transmits the selected instructions to the terminal as audio and visual media. The terminal plays these instructions and displays a visual demonstration on the screen. The user can perform the task according to the instructions by listening to the audio instructions and referring to the visual guide.

[0791] Step 6:

[0792] After completing an operation, the user enters feedback through an information terminal device. The terminal sends this feedback data to a server. The server analyzes the feedback and evaluates whether the problem has been resolved. This analysis involves determining the content of the feedback through text analysis and notifying local support personnel if additional support is needed.

[0793] This series of steps allows users to intuitively and efficiently manage and operate equipment using voice commands.

[0794] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0795] In the system of this invention, not only does the user instruct operations using voice commands, but the system also estimates the user's emotional state by analyzing the voice. This provides more personalized support and improves the user experience.

[0796] Users request support via devices such as smartphones or computers. After the device receives the user's voice command, it sends the voice data to the server. At this time, the server, equipped with an emotion engine, analyzes the voice data and processes it to identify the user's emotions. The emotion engine determines the user's emotions based on characteristics such as the intonation, speed, and pitch of the voice.

[0797] Based on the results of the emotion analysis, the server takes into account the user's emotional state and selects the most appropriate operation guide. For example, if the emotion engine determines that the user is confused, it aims to alleviate the user's anxiety by selecting a guide that provides detailed step-by-step instructions.

[0798] The user guide is tailored to the user's technical skill level, language, and emotional state, and is sent from the server to the terminal. This ensures the guide is tailored to the user's needs, enabling more effective support. Users can follow the instructions while viewing the audio and video displayed on the terminal.

[0799] Furthermore, users can provide feedback on the usefulness of the support, and this feedback is resent to the server via their device. The server analyzes the feedback and, if necessary, compares it with the results of the emotion engine to improve future support methods. Also, if the emotion engine determines from the feedback that the user's emotions remain unstable, it notifies local support staff to prioritize on-site assistance.

[0800] In this way, by utilizing an emotion engine, this system aims to not only solve technical problems but also comprehensively bridge the digital divide by taking into account the psychological aspects of the user.

[0801] The following describes the processing flow.

[0802] Step 1:

[0803] The user inputs a voice command into the device. The device records this using its microphone and temporarily stores it as audio data.

[0804] Step 2:

[0805] The device sends the recorded audio data to the server. This transmission is encrypted to prevent eavesdropping on the audio data.

[0806] Step 3:

[0807] The server receives the audio data and uses a speech recognition engine to convert the audio data into text.

[0808] Step 4:

[0809] The server analyzes the converted text using a natural language processing (NLP) engine to identify the type of support the user is requesting.

[0810] Step 5:

[0811] The server processes the voice data into an emotion engine, analyzing factors such as tone, intensity, and speed to estimate the user's emotional state.

[0812] Step 6:

[0813] The server combines the results of natural language processing and sentiment analysis to select the most appropriate operation guide from the database. This selection takes into account the user's pre-configured technical skill level and emotional state.

[0814] Step 7:

[0815] The server converts the selected operation guide into audio and video instructions and sends them to the terminal. These instructions are designed to be easy for the user to understand.

[0816] Step 8:

[0817] The device provides audio and video instructions to the user. The user can proceed with the operation while watching the on-screen demonstration.

[0818] Step 9:

[0819] Users can provide voice feedback. The device records that feedback and sends it to the server.

[0820] Step 10:

[0821] The server receives feedback and analyzes it using the emotion engine. If the user's emotions remain unstable, it notifies local support staff and prioritizes arranging on-site assistance.

[0822] (Example 2)

[0823] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0824] As modern information and communication technologies advance, users are increasingly required to operate devices based on diverse voice commands. However, conventional systems only perform simple voice recognition without considering the user's emotional state, making it difficult to provide effective support. Furthermore, they lack the flexibility to adapt to the diversity of users' technical abilities and languages. This can potentially expose users to stress and anxiety, potentially impairing their digital technology experience.

[0825] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0826] In this invention, the server includes a computing device means that performs speech recognition and emotion estimation to analyze voice information, a computing device means that selects and presents operational guidance based on the analysis results, and a computing device means that notifies local support personnel and arranges on-site support if the user's problem is not resolved. This enables the provision of personalized operational guides that correspond to the user's emotional state and flexible support that accommodates diversity in technical abilities and languages.

[0827] A "user" refers to an individual who operates the system and issues voice commands.

[0828] "Voice commands" refer to information that users give via voice commands or requests through their devices.

[0829] "Device means" refers to hardware or software functions for receiving voice commands from the user or displaying information.

[0830] "Voice information" refers to the voice data obtained after converting the user's voice commands into digital data.

[0831] "Speech recognition" refers to the technology that analyzes speech information and understands the content of speech commands.

[0832] "Emotion estimation" refers to a technology that determines a user's emotional state by analyzing voice information.

[0833] "Computation device means" refers to a computer system for speech recognition, emotion estimation, and the selection and presentation of operational guidance.

[0834] "Operational guidance" refers to information that provides guidelines and procedures for taking action based on the user's voice commands.

[0835] An "information display device" refers to hardware such as displays and projection equipment used to provide users with instructions on how to operate a system.

[0836] An "image device" refers to a device used to visually present information to a user.

[0837] A "regional support officer" refers to a specialist designated to provide on-site assistance in resolving users' technical problems.

[0838] "On-site support" refers to support activities that involve actually going to the user's location and directly resolving their problems.

[0839] "Multilingual support" refers to a feature that provides assistance in the language each user understands, for users with different language settings.

[0840] "Technical ability" refers to the technical knowledge and operational skills that a user possesses.

[0841] This invention is a system that improves the user experience by allowing users to operate it through voice commands. Users use a terminal such as a smartphone or computer to issue voice commands. The terminal converts the voice into digital data using a microphone or voice input module and transmits it to a server for analysis.

[0842] The server functions as a computing device equipped with advanced emotion estimation capabilities. This server analyzes audio data using speech recognition algorithms to estimate the user's emotional state. It analyzes features such as intonation, speed, and pitch, and utilizes generative AI models to determine the emotional state. For example, it builds emotion estimation models using machine learning frameworks such as TensorFlow and PyTorch.

[0843] Based on the analysis results, the server selects appropriate instructions. These instructions are customized based on the user's emotional state, technical ability, and multilingual support. The server-selected instructions are sent to the terminal, which then provides guidance to the user using a video display or audio output device.

[0844] As a concrete example, suppose a user gives a voice command saying, "I want to set up email on my new device." In this case, the server will assume the user is feeling anxious and select a detailed, step-by-step email setup guide for first-time users, sending it to the device. Alternatively, it can send a prompt to the generative AI model such as: "The user is anxious about setting up email on their new device. Please suggest a gentle and detailed email setup guide." This allows the user to proceed with confidence.

[0845] Furthermore, the server collects user feedback and evaluates the usefulness of voice guidance. Feedback analysis allows for continuous improvement of support services and responses to user requests. By notifying local support staff and arranging on-site assistance as needed, more comprehensive support becomes possible.

[0846] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0847] Step 1:

[0848] The user issues voice commands to the terminal. The input is the user's voice. The terminal uses its microphone to receive the voice as an analog signal. A voice processing module converts the analog signal into digital data, performing noise reduction and sound quality improvement. This digital data is output in the form of voice commands and is ready to be sent to the server.

[0849] Step 2:

[0850] The terminal sends digitized voice commands to the server via a security protocol. The input is digital voice data. The terminal performs data transfer, and the server receives the voice data. The output is the digital voice data transferred to the server.

[0851] Step 3:

[0852] The server analyzes the received digital audio data. The input is audio data. The server uses a speech recognition engine to convert it into text-based voice commands, and then utilizes an emotion estimation engine to determine the user's emotions. It analyzes data such as intonation, speed, and pitch, and outputs an emotion index.

[0853] Step 4:

[0854] The server selects the optimal operation guidance based on voice commands and emotion indicators. The inputs are voice commands and emotion indicators. Using a generative AI model, it infers the most appropriate operation guidance from the voice commands and further adjusts the guidance considering the emotion indicators. The output is customized operation guidance.

[0855] Step 5:

[0856] The server sends the configured operating instructions to the terminal. These instructions include video explanations and audio guides. The input is the customized operating instructions, which are sent to the terminal via data transfer from the server. The output is the operating instructions received by the terminal.

[0857] Step 6:

[0858] The terminal displays operating instructions received from the server to the user. The input is the operating instruction data. The terminal's display and speaker are used to present information to the user visually and audibly. The output is the operating instructions displayed to the user.

[0859] Step 7:

[0860] The user inputs feedback on the usefulness of the operation instructions into the terminal. This input is user feedback information. The terminal collects the input feedback and sends it to the server. The output of this transmission is the feedback data received by the server.

[0861] Step 8:

[0862] The server analyzes the received feedback and considers ways to improve future support. The input is feedback data. It compares this data with sentiment estimation results to evaluate system performance and identify areas for improvement. Furthermore, it sends notifications to local support staff as needed. The output is an improved support plan or support notification.

[0863] (Application Example 2)

[0864] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0865] Current content distribution services are insufficient in suggesting content that is tailored to the individual emotional state of users. Therefore, there is a need for services that are more attentive to users' emotions.

[0866] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0867] In this invention, the server includes means for receiving voice commands from the user and analyzing the voice data to estimate the user's emotional state; means for selecting an operation guide based on the results of voice recognition and emotion analysis and optimizing it according to the user's emotional state; and means for presenting the selected operation guide to the user visually and audibly. This enables personalized content suggestions based on the user's emotional state.

[0868] "Voice commands" refer to a method in which users give specific operational instructions to a system using their voice.

[0869] "Voice data" refers to the digital representation of voice input collected from users.

[0870] "Emotional state" refers to the emotional tendencies and psychological state analyzed from the user's voice.

[0871] "Speech recognition" is a technology that analyzes input speech data and converts it into text.

[0872] An "operation guide" is a set of instructions and procedures provided to a user to operate a system.

[0873] "Presenting information visually and aurally" refers to methods of conveying information to users using screen displays and audio output.

[0874] "Feedback" refers to the act of a user responding to or evaluating the services provided by a system.

[0875] "Emotional analysis" is a process that analyzes voice data to identify the user's emotional state.

[0876] "Improving support services" means optimizing the services and guides provided based on user feedback.

[0877] This invention is a personalized content delivery system that takes into account the user's emotional state, and primarily uses the user's voice input to analyze their emotions and provide optimal content. The system is implemented using the following means.

[0878] The system's terminal receives voice commands from the user. The voice data is sent to the server via the internet. The server converts the voice data into text using the Google Cloud Speech-to-Text API. Then, it uses Azure Text Analytics to perform sentiment analysis on this text data and estimate the user's emotional state.

[0879] Once the user's emotional state has been estimated, the server selects operation guides and content based on the results. The selection criteria include considering the user's psychological feedback and requests such as wanting to relax or become excited. During this process, generative AI models are used to optimize the user interface and operation experience.

[0880] Finally, the selected content and guides are presented to the device visually and aurally. Users can interact with these instructions and provide feedback. This feedback information is then sent back to the server to help improve the quality of future content delivery.

[0881] As a concrete example, consider a scenario where a user watches a movie and says, "Tell me what movie you recommend I watch next." If sentiment analysis of the voice data suggests the user is in a happy mood, a popular local comedy film will be suggested based on this. An example of a prompt would be, "Analyze the voice feedback the user has given on the content they watched and recommend content based on their emotional state."

[0882] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0883] Step 1:

[0884] The terminal receives voice commands from the user. Voice data is input and digitized as an audio file. The received voice data is transferred to a server via the intranet for processing.

[0885] Step 2:

[0886] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. This process transforms the audio data into a parseable text format, making it ready for use in subsequent processes.

[0887] Step 3:

[0888] The server uses Azure Text Analytics to perform sentiment analysis on the converted text data. Here, text data is input, and evaluation metrics indicating the user's emotional state are generated. This quantifies and outputs the user's psychological state.

[0889] Step 4:

[0890] The server selects appropriate content or user guides based on the sentiment analysis results. Here, sentiment evaluation metrics and the user's past history data are used as input, and a generative AI model is employed to select the optimal content. This output forms the basis for the next step.

[0891] Step 5:

[0892] The selected content or operating guide is sent to the device. The content is entered into the device and presented to the user visually and aurally. Through this process, the user receives personalized content.

[0893] Step 6:

[0894] The user provides voice feedback on the content presented by the system. The device then records the user's feedback again as voice data and prepares to send it to the server.

[0895] Step 7:

[0896] The server receives feedback from the user and performs sentiment analysis again. Using the feedback audio data as input, it outputs a new sentiment evaluation metric. This allows the server to confirm user satisfaction and use the results to improve service quality.

[0897] Step 8:

[0898] Based on feedback, the server improves its content selection algorithm and sentiment analysis model as needed. This improvement process uses accumulated sentiment data and feedback as input, enabling the delivery of more accurate content.

[0899] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0900] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0901] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0902] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0903] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0904] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0905] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0906] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0907] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0908] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0909] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0910] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0911] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0912] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0913] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0914] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0915] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0916] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0917] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0918] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0919] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0920] The following is further disclosed regarding the embodiments described above.

[0921] (Claim 1)

[0922] A terminal that receives voice commands from the user,

[0923] A server means that performs speech recognition in order to analyze audio data,

[0924] A server means that selects and provides an operation guide based on the analysis results,

[0925] A terminal device that presents the selected operation guide to the user via audio and video,

[0926] A terminal device that receives user feedback and sends it to a server,

[0927] A server mechanism to notify local support staff and arrange for on-site assistance if the user's problem cannot be resolved.

[0928] Includes system.

[0929] (Claim 2)

[0930] Multilingual support provides a means to offer appropriate language support to users with different language settings.

[0931] The system according to claim 1.

[0932] (Claim 3)

[0933] The system provides means for customizing audio and video instructions according to the user's technical skill level.

[0934] The system according to claim 1.

[0935] "Example 1"

[0936] (Claim 1)

[0937] A computing means for acquiring voice commands from the user,

[0938] A computing device means for performing automatic speech recognition to convert audio data into text information,

[0939] A language analysis means that analyzes textual information and identifies the requested support content,

[0940] A computing device means that selects and provides appropriate operating guidelines based on the analysis results,

[0941] A computing means that presents selected operating guidelines in audio and video,

[0942] A computing means that receives user input and transmits it to the arithmetic unit,

[0943] A computing device means that notifies a technical support person and arranges a visit if the user's problem cannot be resolved.

[0944] Includes system.

[0945] (Claim 2)

[0946] The system incorporates multilingual processing to provide appropriate language support to users with different language settings.

[0947] The system according to claim 1.

[0948] (Claim 3)

[0949] The system provides means for adapting audio and video guidelines based on the user's technical skills.

[0950] The system according to claim 1.

[0951] "Application Example 1"

[0952] (Claim 1)

[0953] Means including an information terminal device that receives voice commands from a user,

[0954] A computer means for performing speech recognition processing to analyze audio data,

[0955] A computing means that selects and provides instructions based on the analysis results,

[0956] Means including an information terminal device that presents the selected instructions to the user audibly and visually,

[0957] Means including an information terminal device that receives user feedback and transmits it to a computer,

[0958] A computing means that notifies local support providers and arranges on-site support if the user's problem cannot be resolved,

[0959] A means for managing safety protection devices based on voice commands,

[0960] A means of providing voice guidance in emergencies

[0961] Includes system.

[0962] (Claim 2)

[0963] The system according to claim 1, comprising means for providing appropriate language support to users with different language settings through multilingual support.

[0964] (Claim 3)

[0965] The system according to claim 1, comprising means for adjusting the provision of voice and visual instructions according to the user's technical ability.

[0966] "Example 2 of combining an emotion engine"

[0967] (Claim 1)

[0968] A device means for acquiring voice commands from the user,

[0969] A computing device means that performs speech recognition and emotion estimation in order to analyze speech information,

[0970] A computing device means that selects and presents operational guidance based on the analysis results,

[0971] Device means for displaying selected operation instructions on an information display device and an image device,

[0972] A device means for obtaining user feedback and transmitting it to a computing device,

[0973] A computing device that notifies local support personnel and arranges on-site support if the user's problem cannot be resolved.

[0974] Includes system.

[0975] (Claim 2)

[0976] Multilingual support provides a means to offer appropriate language support to users with different language settings.

[0977] The system according to claim 1.

[0978] (Claim 3)

[0979] The information display device and image display system are provided with means for making adjustments according to the user's technical capabilities.

[0980] The system according to claim 1.

[0981] "Application example 2 when combining with an emotional engine"

[0982] (Claim 1)

[0983] A means of receiving voice commands from the user, analyzing the voice data, and estimating the user's emotional state,

[0984] A means of selecting an operation guide based on the results of speech recognition and emotion analysis, and optimizing it according to the user's emotional state,

[0985] A means of presenting the selected operation guide to the user visually and aurally,

[0986] A means of receiving user feedback, comparing it with analysis results, and improving support services.

[0987] A system that includes a mechanism for notifying local support staff and arranging on-site assistance if support is insufficient.

[0988] (Claim 2)

[0989] The system according to claim 1, which combines multilingual analysis and sentiment analysis to provide personalized responses to users with different language settings and emotional states.

[0990] (Claim 3)

[0991] The system according to claim 1, comprising means for providing audio and video instructions, which customizes content according to the user's technical skill level and emotional state, and recommends appropriate entertainment content. [Explanation of Symbols]

[0992] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A terminal that receives voice commands from the user, A server means that performs speech recognition in order to analyze audio data, A server means that selects and provides an operation guide based on the analysis results, A terminal device that presents the selected operation guide to the user via audio and video, A terminal device that receives user feedback and sends it to a server, A server mechanism to notify local support staff and arrange for on-site assistance if the user's problem cannot be resolved. Includes system.

2. Multilingual support provides a means to offer appropriate language support to users with different language settings. The system according to claim 1.

3. The system provides means for customizing audio and video instructions according to the user's technical skill level. The system according to claim 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A