System
The system uses video calls and AI to quickly identify and resolve user issues through real-time data analysis, enhancing support efficiency and user satisfaction.
Patent Information
- Application Number
- JP2024123905
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Users face difficulties in identifying and communicating problems effectively, leading to delayed problem resolution and increased resource consumption, and support is often unavailable in multiple languages, affecting user satisfaction.
A system utilizing video calls, generative AI models, and real-time data analysis to identify user problems by capturing and analyzing video and audio data, generating solutions, and providing video tutorials.
Enables efficient problem identification and resolution without detailed user explanation, improving user satisfaction and reducing support costs.
Smart Images

Figure 2026022388000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When users use a service, they are unable to identify the problem and are unable to make an effective inquiry, which delays problem resolution. Supporting such issues also requires significant resources and costs for companies. Furthermore, support is often not available in languages other than Japanese, making it difficult to improve user satisfaction. [Means for solving the problem]
[0005] The present invention provides a system that uses video calls to communicate the status of a problem a user is facing to a system, analyzes the problem using a generative AI model, and provides a solution in real time. Specifically, the system includes: a means for the user to initiate a video call; a means for the terminal to capture and transmit the user's video and audio; a means for the server to analyze the received video and audio data to identify the user's problem; a means for the server to generate a solution to the problem; a means for the server to send an explanation of the generated solution and a video link to the user; and a means for the terminal to display the received explanation of the solution and the video link to the user. The system also includes a means for the server to analyze the GUI status and error messages from the received video data, a means for analyzing the user's comments from the audio data, and a means for identifying the user's problem by integrating the results of the analysis of the video and audio data. The system also includes a means for the server to search a database for an appropriate solution based on the identified problem, a means for generating a video tutorial for solving the problem, a means for transmitting the generated video tutorial to the user, and a means for the terminal to play the received video tutorial. This allows users to effectively communicate their problems and enables companies to efficiently solve problems.
[0006] "User" refers to any person or entity that initiates a video call and uses the Service.
[0007] A "terminal" is a device used by a user, which is equipped with a camera and a microphone and has the function of capturing and transmitting video and audio.
[0008] "Server" refers to the central system that receives, stores, and analyzes video and audio data sent from a user's terminal.
[0009] A "video call" is a communication method that uses a camera and microphone to send and receive images and audio in real time.
[0010] "Video Data" refers to visual information captured by a device's camera.
[0011] "Audio data" refers to sound information captured by the device's microphone.
[0012] A "generative AI model" refers to an artificial intelligence algorithm that analyzes various information such as audio and video and generates appropriate solutions.
[0013] "Analysis" refers to the process by which the server interprets the meaning and status of the information based on the video and audio data it receives.
[0014] "Identifying the problem" refers to the act of identifying the specific obstacle or error the user is encountering.
[0015] "Solution" refers to instructions or steps to resolve an identified problem.
[0016] "Video Link" refers to access to a video that visually explains the solution.
[0017] "Video tutorial" refers to instructional materials in video format that visually illustrate the content of a solution.
[0018] A "database" refers to a collection of information that systematically stores solutions and other necessary information. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The present invention is a system for identifying problems that users encounter when using a service and providing solutions to those problems. This system significantly improves support efficiency because it eliminates the need for users to explain the problem in detail. A specific embodiment of this system is described below.
[0041] Overall flow
[0042] 1. Start a video call
[0043] A user launches a supporting app on their device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone, generates a video call stream, and sends it to the server. The server receives the connection request and creates a new video call session.
[0044] 2. Video and audio capture
[0045] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server, which stores the received video and audio streams in a buffer and prepares them for analysis.
[0046] 3. Data Analysis
[0047] The server extracts each frame from the received video stream and uses an image analysis engine to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine and analyzes what the user is saying. The results of these analyses are combined to identify the problem the user is facing.
[0048] 4. Solution Generation
[0049] The server searches the database for the best solution for the identified problem, generates a video tutorial for solving the problem, creates an explanatory video, and generates a text description of the solution and a video link.
[0050] 5. Providing solutions
[0051] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and if the user clicks on the video, plays the explanatory video.
[0052] Specific examples
[0053] Example 1: If you cannot log in
[0054] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[0055] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[0056] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[0057] Example 2: When an error message appears
[0058] The user uses the camera to capture the situation where an error message is displayed and speaks. The device then sends this to the server.
[0059] The server extracts the specific error message from the video data and analyzes the user's utterances from the audio data. From these results, the server determines that the "specific error message" is the cause.
[0060] The server searches the database for a solution to this error message, generates an explanatory video showing the specific solution method, and sends the generated solution explanation and video link to the user, who then displays it on the device.
[0061] As described above, the present invention provides a system for quickly and accurately identifying problems faced by users and providing appropriate solutions, thereby significantly improving user satisfaction.
[0062] The processing flow will be explained below.
[0063] Step 1:
[0064] User: Launch the support app and click the "Start Video Call" button.
[0065] Step 2:
[0066] Device: Asks the user for permission to use the camera and microphone. If permission is granted, turns on the camera and microphone and generates the video call stream.
[0067] Step 3:
[0068] Terminal: Sends the video call stream to the server. Video and audio data are encoded in real time and continuously transmitted to the server.
[0069] Step 4:
[0070] Server: Receives connection requests from users' devices and creates new video call sessions. It buffers the received video and audio streams.
[0071] Step 5:
[0072] Server: Extracts each frame from the received video stream and uses an image analysis engine to identify GUI states and error messages.
[0073] Step 6:
[0074] Server: Converts the audio stream into text using a speech analysis engine. Analyzes what the user says and extracts keywords.
[0075] Step 7:
[0076] Server: Combines the results of video and audio analysis, applies rule-based filters, and identifies the problem the user is facing.
[0077] Step 8:
[0078] Server: Searches for a suitable solution from a database based on the identified problem, and generates an instructional video demonstrating the solution if necessary.
[0079] Step 9:
[0080] Server: Generates a text description of the solution and a video link and sends them to the user's device.
[0081] Step 10:
[0082] On the device: The received text description and video link are displayed to the user. If the user clicks on the video link, the video player is launched and the explanatory video is played.
[0083] Step 11:
[0084] User: Check the solution displayed on the device and take the necessary action. If the solution is insufficient, ask for additional support via video call again.
[0085] Example 1
[0086] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0087] Conventional user support systems require users to describe their problems in detail, a process that takes time and effort. Furthermore, identifying problems and providing solutions is often done manually, which is often inefficient. These issues result in lower user satisfaction and increased support costs.
[0088] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0089] In this invention, the server includes means for analyzing the status of the graphical user interface and error messages from the received video data, means for analyzing the content of user comments from the received audio data, and means for identifying the user's problem by integrating the results of the analysis of the video data and the audio data, thereby making it possible to efficiently and quickly identify the problem and provide a solution without the user having to explain the problem in detail.
[0090] The "means for initiating a video call" refers to the operation performed by a user to initiate a video call using an application on a terminal, and the combination of software and hardware that realizes this operation.
[0091] The "means for capturing and transmitting video and audio" is a mechanism for capturing video and audio data in real time using the camera and microphone of the terminal, encoding this data, and transmitting it to the server.
[0092] The "means for analyzing the received video and audio data to identify the problem the user is facing" is a software process in which the server processes the received video and audio data using an analysis engine to identify the problem or issue the user is facing.
[0093] The "means for generating a solution" is a combination of software and hardware that allows the server to search for an appropriate solution from the database and generate a solution to provide to the user based on the searched solution.
[0094] The "means for transmitting the generated solution description and video link to the user's terminal" is a communication means for transmitting the server-generated text description of the solution and the associated video link to the user's terminal via a network.
[0095] The "means for displaying the received solution description and video link to the user" is a mechanism including a user interface for displaying the text description and video link of the solution received by the terminal from the server in a format that is easy for the user to view.
[0096] "Means for analyzing the status of the graphical user interface and error messages" refers to a software process in which the server analyzes video data to extract information such as on-screen buttons, menus, and error messages, and then uses this information to identify the user's operational status and problems.
[0097] "Means for analyzing the user's speech" refers to a software process by which the server converts the speech data into text using speech recognition technology and understands from the text what issue the user is reporting.
[0098] The "means for searching for an appropriate solution from the database" is a mechanism by which the server searches for and obtains the most appropriate solution for a specified problem from among the solutions stored in the database.
[0099] The "means for generating video tutorials" refers to video generation technology and software that allows the server to create videos that include specific steps to solve a user's problem.
[0100] The "means for transmitting to the user's terminal" refers to a communication means for transmitting the video tutorial and solution information generated by the server to the user's terminal via a network.
[0101] The "means for playing the video tutorial to the user" refers to a mechanism including a media player that allows the user to play and watch the video tutorial received by the user's terminal.
[0102] MODE FOR CARRYING OUT THE INVENTION
[0103] The present invention is a system for identifying problems users encounter when using a service and providing solutions to those problems. This system significantly improves support efficiency because it does not require users to describe the problem in detail.
[0104] Starting a video call
[0105] When a user launches a supporting app and clicks the "Start Video Call" button, the device requests permission to use the camera and microphone. If permission is granted, the device activates the camera and microphone, generates a video call stream, and sends it to the server. This stream uses technologies such as WebRTC. The server receives the connection request and creates a new video call session.
[0106] Video and audio capture
[0107] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server. The encoding is done using H.264 and AAC encoders. The server stores the received video and audio streams in a buffer and prepares them for analysis.
[0108] Data analysis
[0109] The server extracts each frame from the received video stream and uses an image analysis engine (e.g., OpenCV or TensorFlow) to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine (e.g., Google Speech-to-Text or IBM Watson Speech to Text) and analyzes what the user is saying. The results of this analysis are combined to identify the problem the user is facing.
[0110] Solution Generation
[0111] The server searches for the best solution for the identified problem from a database (e.g., MySQL or MongoDB). It then generates a video tutorial to solve the problem. Adobe After Effects and FFmpeg are used to generate the video. It then generates a text description of the solution and a video link.
[0112] Providing solutions
[0113] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user. When the user clicks the video link, the device plays the explanatory video.
[0114] Specific examples
[0115] Example 1: If you cannot log in
[0116] The user launches the support app and says, "I can't log in." The device uses its camera and microphone to send video and audio data to the server. The server identifies the blank fields on the login screen from the video data and analyzes the audio data. This allows the server to identify the problem: "I can't log in because the password field is blank." The server searches its database for a solution to the correct login method and generates a video tutorial showing the login procedure along with the message, "An entry is required in the password field." This is sent to the user's device, which then displays an explanation of the solution and a video link.
[0117] Example 2: When an error message appears
[0118] The user films the situation in which "an error message is displayed" and speaks about it. The device sends this to the server. The server extracts the specific error message from the video data and analyzes the user's statement from the audio data. As a result, the server determines that the "specific error message" is the cause. The server searches a database for a solution to this error message and generates an explanatory video showing a specific method of dealing with the problem. The generated solution explanation and video link are sent to the user, and the device displays them to the user.
[0119] Example prompts for generative AI models
[0120] Below is an example of a prompt sentence to input to the generative AI model.
[0121] 1. "A user initiates a video call. Video and audio are captured and sent to the server."
[0122] 2. "The server analyzes the video data received and checks for error messages and the status of the GUI."
[0123] 3. "Based on the identified problem, we generate optimal solutions and tutorial videos."
[0124] As described above, this system is able to quickly and accurately identify user problems and provide appropriate solutions, which is expected to improve user satisfaction and significantly improve support efficiency.
[0125] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0126] Step 1:
[0127] Starting a video call
[0128] The user launches the support app and clicks the "Start Video Call" button.
[0129] Input: User action.
[0130] The device will display a popup requesting permission to use the camera and microphone, and if the user allows it, the camera and microphone will be activated.
[0131] Input: User permission.
[0132] The device generates a video call stream and sends it to the server using WebRTC technology. The device sets the stream's protocol information and starts the stream.
[0133] Output: Video call stream.
[0134] The server receives the connection request and creates a new video call session. The server generates a session ID and sends a confirmation message to the device.
[0135] Output: A new video call session.
[0136] Step 2:
[0137] Video and audio capture
[0138] The device uses a camera to capture video data and a microphone to capture audio data.
[0139] Input: Video and audio data.
[0140] The captured video and audio data is encoded in real time using H.264 and AAC encoders.
[0141] The terminal transmits the encoded video and audio data to the server.
[0142] Output: Encoded video and audio data.
[0143] The server stores the received data in a buffer and prepares it for analysis. The server manages the buffer memory and checks the data integrity.
[0144] Output: The data stored in the buffer.
[0145] Step 3:
[0146] Data analysis
[0147] The server extracts each frame from the received video stream using a video analysis engine (e.g., OpenCV).
[0148] Input: Buffered video data.
[0149] The server uses an image analysis engine to identify GUI states (e.g., buttons, input fields, error messages).
[0150] Input: Video data for each frame.
[0151] Output: Data of the identified GUI element.
[0152] The server converts the audio stream into text using a speech analysis engine (e.g., Google Speech-to-Text), which is then analyzed by a natural language processing engine.
[0153] Input: Buffered audio data.
[0154] Output: The audio data converted to text.
[0155] The server combines the results of the video and audio analysis to identify the problem the user is facing.
[0156] Input: Parsed GUI data and text data.
[0157] Output: Identified issues.
[0158] Step 4:
[0159] Solution Generation
[0160] The server searches for a solution to the identified problem from a database that stores information about existing solutions.
[0161] Input: Identified problem.
[0162] Output: The solution found.
[0163] The server uses a video generation engine (e.g. Adobe After Effects) to generate a video tutorial for solving the problem. The video is automatically generated based on the steps you specify.
[0164] Input: The solution found.
[0165] Output: The generated video tutorial.
[0166] The server generates a message containing a text description of the solution and a video link.
[0167] Input: Retrieved solutions and generated video tutorials.
[0168] Output: Solution description and video link.
[0169] Step 5:
[0170] Providing solutions
[0171] The server generates a solution explanation and sends a video link to the user's device.
[0172] Input: Solution description and video link.
[0173] Output: Data sent to the user's terminal.
[0174] The device displays the received solution explanation and video link to the user.
[0175] Input: Data received from the server.
[0176] Output: The solution description and video link that is displayed to the user.
[0177] When the user clicks on the video link, the device will play the instructional video.
[0178] Input: Click action on the video link.
[0179] Output: The instruction video that will be played.
[0180] (Application example 1)
[0181] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0182] When a machine or robot operating in a factory begins to operate abnormally, it is necessary to quickly and accurately identify the problem and provide an appropriate solution. However, currently, it takes a lot of time for operators to identify the problem and find a solution. In addition, finding an appropriate solution requires specialized knowledge, which leads to a decrease in production efficiency. Therefore, there is a need to introduce a system that can detect abnormalities in real time, efficiently identify the problem, and provide a solution.
[0183] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0184] In this invention, the server includes a means for detecting abnormal operation of the machine, a means for transmitting data captured by the machine to the server, and a means for identifying an optimal solution for the machine and instructing the machine on how to deal with the problem. This makes it possible to quickly detect abnormal operation, efficiently identify the problem, and provide an appropriate solution.
[0185] "Means for a user to initiate a video call" refers to software or hardware functionality that allows a user to initiate a video call session to receive remote support.
[0186] "Means for the terminal to capture and transmit the user's video and audio" refers to a function that enables the terminal to use a camera and microphone to capture the user's video and audio and transmit them to a server.
[0187] "Means for analyzing the video and audio data received by the server to identify the problem the user is facing" is a function for analyzing the video and audio data received by the server and identifying the problem the user is facing.
[0188] The "means by which the server generates a solution to the identified problem" is a function by which the server generates an appropriate solution to the problem after identifying the problem.
[0189] "Means for transmitting the explanation of the solution generated by the server and a video link to the user's terminal" is a function for transmitting the explanation of the solution generated by the server and a related video link to the user's terminal.
[0190] The "means for displaying to the user the explanation of the solution and the video link received by the terminal" is a function for displaying to the user the explanation of the solution and the video link received by the user's terminal.
[0191] "Means for machines to detect abnormal operation" refers to software or hardware functions that allow machines operating in a factory to monitor their operating status and detect any abnormalities that may occur.
[0192] "Means for transmitting data captured by the machine to the server" is a function for transmitting video and audio data captured by the machine to the server when an abnormality is detected.
[0193] "Means for identifying the optimal solution for the machine and instructing the machine on how to respond" is a function that enables the server to analyze a problem, identify an appropriate solution, and instruct the machine on how to respond.
[0194] MODE FOR CARRYING OUT THE INVENTION
[0195] A system for implementing this invention includes various means for quickly and accurately identifying a user's problem and providing an appropriate solution. Specifically, a user initiates a video call, and the video and audio data captured by the terminal is sent to a server, which analyzes the data and identifies the problem. The server then generates a solution and provides it to the user. In a factory, if a machine detects an abnormal operation, it notifies the server, which then analyzes the situation and provides an optimal solution.
[0196] Hardware and Software
[0197] Factory robots: equipped with cameras, microphones, and communication modules, including sensors to detect abnormal behavior.
[0198] Server: A video analysis engine (e.g., OpenCV), an audio analysis engine (e.g., Google Speech-to-Text API), a database (e.g., MySQL), and a solution generation module (Python script).
[0199] Device: Includes the device used by the user (smartphone, tablet, PC, etc.).
[0200] Data processing and calculation
[0201] The server first receives video and audio data sent from terminals and factory robots. The received video data is analyzed using an image analysis engine (OpenCV) to identify GUI status and error messages. Meanwhile, the audio data is converted into text using a speech analysis engine (Google Speech-to-Text API), and the user's speech is analyzed.
[0202] The analyzed data is integrated and the server uses this information to identify the problems the user is facing. For the identified problems, the server searches the database (MySQL) for the best solution and generates a text explanation and a video tutorial as the solution.
[0203] The generated solution is sent from the server to the user's device, where it is displayed to the user, along with notifications and explanatory video links sent from the device.
[0204] Specific examples
[0205] 1. Resolving login issues:
[0206] The user says, "I can't log in."
[0207] The device uses a camera and microphone to capture video and audio and transmits them to a server.
[0208] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[0209] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays the solution explanation and video link to the user.
[0210] Example prompt sentence:
[0211] Identify issues from video call data and provide solutions.
[0212] Audio data: "I can't log in"
[0213] Video data: Screenshot of the login screen
[0214] 2. Machine Anomaly Detection and Resolution:
[0215] A factory machine detects abnormal operation.
[0216] The machine uses a camera and microphone to capture video and audio and transmits it to a server.
[0217] The server identifies specific error messages and abnormal behavior patterns from the video data and also analyzes information from the audio data. Based on the results, the server identifies the cause of the abnormality and searches a database for an appropriate solution.
[0218] The server instructs the machine on the generated solutions and workarounds, allowing it to make the necessary corrections.
[0219] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0220] Step 1:
[0221] The user launches the supporting app on the device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone. The input obtained here is the user's video and audio, which are generated as a video call stream and sent to the server.
[0222] Step 2:
[0223] The server buffers the video call stream received from the device. The received video data is divided into frames, and the audio data is encoded. The input to this analysis preparation stage is the video and audio streams, and the output is data ready for analysis.
[0224] Step 3:
[0225] The server uses an image analysis engine (OpenCV) to analyze each frame of the received video stream. Specifically, it performs processing to identify the GUI status and error messages. The input is the received video data, and the output is information about the GUI status and error messages.
[0226] Step 4:
[0227] The server converts the received audio stream into text using a speech analysis engine (Google Speech-to-Text API). This converts the user's input voice data into text data and analyzes the user's speech. The input is voice data, and the output is text data.
[0228] Step 5:
[0229] The server integrates the results of the video data analysis obtained in step 3 and the results of the audio data analysis obtained in step 4. This identifies the problem the user is facing. The inputs are the video analysis results and the audio analysis results, and the output is information about the identified problem.
[0230] Step 6:
[0231] The server searches for a suitable solution from a database (MySQL) based on the identified problem. In this step, the input is information about the identified problem and the output is the solution retrieved from the database.
[0232] Step 7:
[0233] The server generates a video tutorial for solving the problem based on the solution to the identified problem. The generated solution includes a text explanation and a video link. The input is the searched solution, and the output is the video tutorial and its link.
[0234] Step 8:
[0235] The server sends the generated solution explanation and video link to the user's device. The input is the video tutorial and link, and the output is the data sent to the user's device.
[0236] Step 9:
[0237] The device displays the received solution explanation and video link to the user. Specifically, when the user clicks on the video link, the explanatory video is played. The input is the received data, and the output is the solution explanation and video link displayed to the user.
[0238] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0239] The present invention is a system that identifies problems users encounter when using services, taking into account the user's emotional state, and provides solutions. Because the user does not need to explain the problem in detail, support efficiency can be significantly improved. A specific embodiment of this system is described below.
[0240] Overall flow
[0241] 1. Start a video call
[0242] A user launches a supporting app on their device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone, generates a video call stream, and sends it to the server. The server receives the connection request and creates a new video call session.
[0243] 2. Video and audio capture
[0244] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server, which stores the received video and audio streams in a buffer and prepares them for analysis.
[0245] 3. Data Analysis
[0246] The server extracts each frame from the received video stream and uses an image analysis engine to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine and analyzes what the user is saying. The results of these analyses are combined to identify the problem the user is facing.
[0247] 4. Emotion analysis
[0248] The server uses an emotion engine that recognizes the user's emotions from the video and audio data, for example, facial expressions and tone of voice to determine whether the user is feeling impatient, angry, sad, etc.
[0249] 5. Solution Generation
[0250] The server searches for the best solution from the database based on the identified problem and the user's emotional state. If necessary, it adjusts the priority and presentation of the solutions taking the emotional state into account. It also generates an appropriate video tutorial for solving the problem and creates an explanatory video. It also generates a text description and video link for the solution.
[0251] 6. Providing solutions
[0252] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and if the user clicks on the video, the device launches a video player and plays the explanatory video.
[0253] Specific examples
[0254] Example 1: If you cannot log in
[0255] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[0256] Emotion analysis: The server determines from video and audio data that the user is feeling anxious, and determines that a solution needs to be offered quickly.
[0257] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[0258] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[0259] Example 2: When an error message appears
[0260] The user uses the camera to capture the situation where an error message is displayed and speaks. The device then sends this to the server.
[0261] Sentiment analysis: The server determines from video and audio data that the user is angry, and determines that a solution should be provided in a calm and easy-to-understand format.
[0262] The server extracts the specific error message from the video data and analyzes the user's utterances from the audio data. From these results, the server determines that the "specific error message" is the cause.
[0263] The server searches the database for a solution to this error message, generates an explanatory video showing the specific solution method, and sends the generated solution explanation and video link to the user, who then displays it on the device.
[0264] As described above, the present invention provides a system that quickly and accurately identifies the problems a user is facing, taking into account the user's emotional state, and provides appropriate solutions, thereby significantly improving user satisfaction.
[0265] The processing flow will be explained below.
[0266] Step 1:
[0267] User: Launch the support app and click the "Start Video Call" button.
[0268] Step 2:
[0269] Device: Asks the user for permission to use the camera and microphone. If permission is granted, turns on the camera and microphone.
[0270] Step 3:
[0271] Terminal: Generates the video call stream, encodes the video and audio data in real time, and sends it to the server.
[0272] Step 4:
[0273] Server: Receives connection requests from users' devices and creates new video call sessions. It buffers the received video and audio streams.
[0274] Step 5:
[0275] Server: Extracts each frame from the video stream and uses an image analysis engine to identify GUI states and error messages.
[0276] Step 6:
[0277] Server: The audio stream is converted into text using a speech analysis engine, and the content of what the user is saying is analyzed.
[0278] Step 7:
[0279] Server: Integrates the results of video and audio analysis and applies rule-based filters to identify the issues the user is facing.
[0280] Step 8:
[0281] Server: Uses an emotion engine to recognize the user's emotions from video and audio data. Emotions are analyzed from the user's facial expressions and tone of voice.
[0282] Step 9:
[0283] Server: Searches for an appropriate solution from a database based on the identified problem and the user's emotional state.
[0284] Step 10:
[0285] Server: Tailors solutions to user emotions and prioritizes them as needed. Generates problem-solving video tutorials and creates instructional videos.
[0286] Step 11:
[0287] Server: Generates a text description of the solution and a video link and sends them to the user's device.
[0288] Step 12:
[0289] On your device: Display the received text description and video link to the user.
[0290] Step 13:
[0291] User: Clicks on the video link and watches a video that explains how to solve the problem.
[0292] Step 14:
[0293] Device: Plays instructional videos and helps users solve problems by following instructions.
[0294] The above steps create a system that can quickly and accurately resolve problems users encounter while using the service. The introduction of an emotion engine makes it possible to provide optimal support based on the user's emotional state.
[0295] Example 2
[0296] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0297] In today's world, users face a wide variety of problems when using services. Support for resolving these problems typically requires users to explain the problem in detail. However, if users are unable to clearly express their problems, solving the problem increases the time and effort required. Furthermore, presenting solutions without considering the user's emotional state can detract from the user experience. To address these challenges, a system is needed that can quickly and accurately identify problems and provide optimal solutions, taking the user's emotions into account.
[0298] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0299] In this invention, the server includes means for analyzing the user's emotions from the received video data and audio data, means for generating an appropriate solution based on the identified problem and the user's emotional state, and means for transmitting an explanation of the generated solution and a video link to the user's terminal. This makes it possible to quickly and accurately identify the problem and provide an optimal solution while taking into account the user's emotions, without the user having to clearly explain the problem.
[0300] "User" means a person who uses the System to request support for the Service.
[0301] "Video call" refers to a real-time means of video and audio communication between a user and a support system.
[0302] A "terminal" is a device that a user uses to interact with the support system, such as a smartphone or computer.
[0303] A "camera" is a device for capturing images that is built into or attached to a terminal.
[0304] A "microphone" is a device for capturing sound that is built into a terminal or attached externally.
[0305] "Server" is a central processing system that analyzes data received through the video call and provides appropriate solutions to the user.
[0306] "Video data" refers to image information captured by a camera.
[0307] "Audio Data" means audio information captured by a microphone.
[0308] "Analysis" is the act of extracting and processing information based on video and audio data.
[0309] "Emotion analysis" is the process of identifying a user's emotional state based on information such as facial expressions and tone of voice.
[0310] A "solution" is a specific procedure or method provided to solve a problem a user is facing.
[0311] A "video link" is a hyperlink that indicates where to watch an explanatory video, etc.
[0312] "Display" is the act of visualizing information such as text or video links on a device screen.
[0313] A "video player" is software or an application that plays an explanatory video when you click on a video link.
[0314] MODE FOR CARRYING OUT THE INVENTION
[0315] The present invention is a system for identifying problems users encounter when using services, taking into account the user's emotional state, and providing solutions, allowing users to receive prompt and accurate support without having to explain their problems in detail.
[0316] Hardware and software used
[0317] Hardware
[0318] Device: The smartphone or computer used by the user.
[0319] Camera: A built-in or external video capture device.
[0320] Microphone: Built-in or external audio capture device.
[0321] Server: A central processing system that processes and analyzes data.
[0322] software
[0323] Supporting App: An application that allows users to initiate video calls.
[0324] Image analysis engine: For example, OpenCV is used to analyze video data.
[0325] Speech analysis engine: Converts voice data into text using, for example, the Google Speech-to-Text API.
[0326] Sentiment analysis engine: Analyze the user's emotional state using, for example, the Affectiva SDK.
[0327] System action
[0328] Starting a video call
[0329] 1. The user launches the supporting app on their device and clicks the "Start Video Call" button.
[0330] 2. The device requests permission to use the camera and microphone, and if the user grants permission, it activates the camera and microphone, generates a video call stream, and sends it to the server.
[0331] 3. The server receives the connection request and creates a new video call session.
[0332] Video and audio capture
[0333] 1. The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server.
[0334] 2. The server buffers the received video and audio streams and prepares them for analysis.
[0335] Data analysis
[0336] 1. The server extracts each frame from the received video stream and uses an image analysis engine to identify the status of the graphical user interface and error messages.
[0337] 2. At the same time, the audio stream is converted into text by a speech analysis engine and the content of what the user is saying is analyzed.
[0338] 3. Synthesize the analysis results to identify the problems users are facing.
[0339] Emotion analysis
[0340] 1. The server uses an emotion analysis engine to recognize the user's emotions from video and audio data.
[0341] 2. For example, by looking at facial expressions and tone of voice, it can identify whether the user is feeling anxious, angry, sad, or other emotions.
[0342] Solution Generation
[0343] 1. The server searches the database for the best solution based on the identified problem and the user's emotional state.
[0344] 2. Adjust the priorities and presentation of solutions to take into account emotional states.
[0345] 3. Generate appropriate video tutorials and create instructional videos to solve problems.
[0346] 4. Create a text description and video link of the generated solution.
[0347] Providing solutions
[0348] 1. The server sends the generated solution explanation and video link to the user's device.
[0349] 2. The device displays the received text description and video link to the user.
[0350] 3. When the user clicks on the video link, the device launches a video player and plays the explanatory video.
[0351] Specific examples
[0352] Example 1: If you cannot log in
[0353] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[0354] Emotion analysis: The server determines from video and audio data that the user is feeling anxious, and determines that a solution needs to be offered quickly.
[0355] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data, thereby identifying the problem of "I can't log in because the password field is blank."
[0356] The server searches the database for a solution to the correct login method, generates a message saying "Password field must be filled in" and a video tutorial showing the login procedure, and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[0357] Prompt Sentence Examples
[0358] "Users report issues via video call, you analyze them including sentiment analysis, and provide solutions with video tutorials. Users say 'I can't log in.'"
[0359] The above is an example of the present invention. This system enables the user to quickly and accurately identify problems and provide optimal solutions, taking into account emotions, without the user having to clearly explain the problem.
[0360] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0361] Specific explanation of processing steps
[0362] Step 1:
[0363] Starting a video call
[0364] Input: The user launches a supporting app and clicks the "Start Video Call" button.
[0365] How it works: A user launches a supporting app and clicks the button for a video call.
[0366] Output: The device displays a dialog to allow use of the camera and microphone.
[0367] Example of how it works:
[0368] When the user clicks the "Start Video Call" button, a dialog box will appear requesting permission to use the device's camera and microphone. If the user clicks "Allow," the camera and microphone will be activated.
[0369] Step 2:
[0370] Video and audio capture
[0371] Input: Camera and microphone activated.
[0372] How it works: The device captures video with its camera and audio with its microphone.
[0373] Output: Encoded video and audio data is generated and sent to the server.
[0374] Example of how it works:
[0375] The user's face and voice are captured in real time by a camera and microphone, encoded, and then streamed to a server.
[0376] Step 3:
[0377] Receiving and buffering data
[0378] Input: Video and audio data sent to the server.
[0379] Operation: The server stores the received video and audio data in a temporary buffer and prepares it for analysis.
[0380] Output: The data stored in the buffer is passed to the analysis engine.
[0381] Example of how it works:
[0382] When the server receives the video and audio streaming data, it is temporarily stored in a buffer, through which the data is passed to the next analysis step.
[0383] Step 4:
[0384] Data analysis
[0385] Input: Buffered video and audio data.
[0386] How it works: The server uses an image analysis engine (e.g., OpenCV) to analyze each frame of video data and recognize graphical user interfaces and error messages, while simultaneously using a speech analysis engine (e.g., Google Speech-to-Text API) to convert the audio data into text and analyze what the user is saying.
[0387] Output: Text data of the identified problem and what the user said.
[0388] Example of how it works:
[0389] A specific error message is read from the video data, and the statement "I can't log in" is converted into text from the audio data.
[0390] Step 5:
[0391] Emotion analysis
[0392] Input: Buffered video and audio data, and the identified issue.
[0393] How it works: The server uses an emotion analysis engine (e.g., Affectiva SDK) to analyze the user's emotional state, for example, by identifying the user's emotions (annoyance, anger, sadness, etc.) from facial expressions and tone of voice.
[0394] Output: Analysis data of the user's emotional state.
[0395] Example of how it works:
[0396] A video of the user's impatient face after seeing the identified error message is analyzed to identify the emotion of impatience.
[0397] Step 6:
[0398] Solution Generation
[0399] Input: Analysis data of identified problems and emotional states.
[0400] How it works: The server searches for the best solution from a database based on the identified problem and the user's emotional state, adjusts the priority and presentation of solutions as needed, and generates a video tutorial for solving the problem and creates an instructional video.
[0401] Output: A text description of the solution and a video link.
[0402] Example of how it works:
[0403] The server generates a "Password field is required" message and a video tutorial showing login steps.
[0404] Step 7:
[0405] Providing solutions
[0406] Input: A text description of the solution and a video link.
[0407] Operation: The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and when the user clicks the video link, it launches a video player and plays the explanatory video.
[0408] Output: The solution shown to the user and a video link.
[0409] Example of how it works:
[0410] The device screen displays the message "Password field requires input" along with a video link. When the user clicks the link, a video player launches and a video with the solution is played.
[0411] (Application example 2)
[0412] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0413] When users encounter security-related problems, conventional support systems often delay problem identification and solution provision, making it difficult to provide a prompt and appropriate response to the user. Furthermore, users who are in an emotional state, such as frustration, anger, or sadness, may find it difficult to understand the support content. This tends to reduce user satisfaction and prolong the problem. Therefore, there is a need for a system that not only speeds up problem identification but also provides solutions that take into account the user's emotional state.
[0414] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing received video and audio data to identify a problem faced by the user, means for analyzing the user's emotional state from the received video and audio data, and means for generating a solution based on the identified problem and the user's emotional state. This makes it possible, when a user faces a security-related problem, to quickly identify the problem and provide an effective solution that takes the user's emotional state into consideration.
[0415] "Video calling" is a function that allows users to communicate remotely in real time using audio and video.
[0416] A "terminal" is a device that a user uses to initiate a video call and capture video and audio. Examples include a smartphone or tablet.
[0417] "Server" is a central processing unit that analyzes video and audio data received from users, identifies problems, and provides solutions.
[0418] "Video data" refers to video information captured by a camera, including the user's face and the state of the screen interface.
[0419] "Audio data" is acoustic information captured by a microphone, including the user's voice and what is being said.
[0420] "Analysis" is the process by which the server extracts useful information from the data it receives and understands its meaning.
[0421] "Emotional state" refers to the emotions that a user expresses through video and audio data, such as impatience, anger, sadness, etc.
[0422] A "solution" is a method, procedure, or supporting information provided by the server to solve an identified problem.
[0423] "Description and Video Link" is a textual summary of the solution and access to a tutorial video related to that solution.
[0424] "Database" means an information management system that accumulates and stores solutions and related information in a searchable format.
[0425] "Video tutorials" are video contents that visually explain to users how to solve problems.
[0426] Overall flow
[0427] The invention includes a system for quickly and effectively responding when a user experiences a security-related problem, which takes into account the user's emotional state to identify the problem and provide a solution.
[0428] Hardware and Software
[0429] Hardware used:
[0430] Smartphone or tablet (e.g. iPhone, Android device)
[0431] Server (high-performance server for sensing and analysis, such as Amazon Web Services)
[0432] Software used:
[0433] Video calling libraries (e.g. WebRTC)
[0434] Sentiment analysis engine (e.g. Microsoft Azure Emotion API)
[0435] Speech analysis engine (e.g. Google Cloud Speech-to-Text)
[0436] A database (e.g., MySQL or PostgreSQL)
[0437] Video platform (e.g. YouTube API)
[0438] System Operation
[0439] 1. Start a video call
[0440] A user launches a specific supporting application on a smartphone or tablet. This application uses WebRTC to provide video calling functionality. When the user clicks the "Start Video Call" button, the camera and microphone are activated and video and audio capture begins.
[0441] 2. Capture of video and audio data
[0442] Video data captured by the smartphone camera and audio data captured by the microphone are encoded in real time and sent to a server, which stores the data in a buffer and prepares it for analysis.
[0443] 3. Data Analysis
[0444] The server divides the received video stream into frames and identifies the face and screen interface state. Libraries such as OpenCV are used for image analysis. Additionally, audio data is converted into text using Google Cloud Speech-to-Text, allowing the user's speech to be analyzed.
[0445] 4. Emotion analysis
[0446] The server uses the received video and audio data to determine the user's emotional state. It uses the Microsoft Azure Emotion API to analyze emotions from facial expressions and complements emotions from the user's tone of voice.
[0447] 5. Solution Generation
[0448] The server searches for an appropriate solution from a database based on the identified problem and the user's emotional state, and then generates a problem-solving explanation text and related tutorial video using the YouTube API.
[0449] 6. Providing solutions
[0450] The server sends the generated solution explanation and video link to the user's device, which displays it to the user. When the user clicks on the video link, a video player is launched and the explanatory video is played.
[0451] Specific examples
[0452] Sentiment analysis prompt example:
[0453] Identify the user's emotions from their facial expressions and tone of voice, specifically whether they are anger, sadness, confusion, or impatience.
[0454] Solution search prompt example:
[0455] Find a solution to the security issue based on the problem the user is facing and their emotional state. Describe your recommended solution and steps, and provide links to relevant video tutorials.
[0456] Thus, an embodiment of the invention provides a system that quickly and accurately solves security problems while taking into account the emotional state of the user.
[0457] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0458] Step 1:
[0459] A user launches a supporting application and clicks the "Start Video Call" button. This receives user input and activates the camera and microphone as output. This prepares the application for capturing video and audio. Specifically, the application uses the WebRTC library to request permission to use the camera and microphone.
[0460] Step 2:
[0461] The device begins capturing video data with the camera and audio data with the microphone. The captured video and audio input data is encoded and sent to the server in real time. The specific operation is a process that uses WebRTC to generate video and audio streams and send the data to the server. The output is the video and audio data sent to the server.
[0462] Step 3:
[0463] The server splits the video stream into frames and uses an image analysis engine (e.g., OpenCV) to determine the GUI state and error messages. This is the process of analyzing the captured video data (input) and extracting specific screen elements and messages (output). The specific operation is to apply image processing algorithms and identify important elements.
[0464] Step 4:
[0465] The server converts the audio stream into text using Google Cloud Speech-to-Text. The input is the transmitted audio data, and the output is the analyzed text data. The specific operation is to send the audio data to the speech analysis engine and obtain the resulting text.
[0466] Step 5:
[0467] The server integrates the analysis results of the video and audio data to identify the problem the user is facing. This is the process of combining the outputs from the image analysis engine and the audio analysis engine to identify the problem. Specifically, it centrally manages the analysis results and applies an algorithm to identify the cause of the problem. The input is the image analysis results and the audio analysis results, and the output is the identified problem.
[0468] Step 6:
[0469] The server performs emotion analysis from the video and audio data. It uses the Microsoft Azure Emotion API to identify the user's emotional state. At this stage, the input is the user's video and audio data, and the output is the analyzed emotional state. The specific operation is to send the data to the emotion analysis engine and obtain the resulting emotional information.
[0470] Step 7:
[0471] The server searches for an appropriate solution from a database based on the identified problem and the user's emotional state. The input is the analysis results and the user's emotional state, and the output is an appropriate solution. The specific operation is to quickly extract a solution using a database search algorithm.
[0472] Step 8:
[0473] The server generates an explanatory video to solve the problem. It also generates links to related tutorial videos using the YouTube API. The input is the searched solution, and the output is a solution explanation including a video link. The specific operation is the process of linking with the YouTube API and extracting the appropriate video.
[0474] Step 9:
[0475] The server sends the generated solution explanation and video link to the user's device. The input is the solution explanation and video link, and the output is data transmission to the user's device. The specific operation is to send data to the user's device via network communication.
[0476] Step 10:
[0477] The device displays the received solution explanation and video link to the user. When the user clicks on the video link, a video player on the device is launched and the explanatory video is played. The input is the solution explanation and video link sent from the server, and the output is the display and video playback on the device. The specific operation is to display the solution on the user interface and launch the video player.
[0478] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0479] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0480] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0481] [Second embodiment]
[0482] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0483] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0484] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0485] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0486] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0487] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0488] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0489] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0490] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0491] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0492] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0493] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0494] The present invention is a system for identifying problems that users encounter when using a service and providing solutions to those problems. This system significantly improves support efficiency because it eliminates the need for users to explain the problem in detail. A specific embodiment of this system is described below.
[0495] Overall flow
[0496] 1. Start a video call
[0497] A user launches a supporting app on their device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone, generates a video call stream, and sends it to the server. The server receives the connection request and creates a new video call session.
[0498] 2. Video and audio capture
[0499] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server, which stores the received video and audio streams in a buffer and prepares them for analysis.
[0500] 3. Data Analysis
[0501] The server extracts each frame from the received video stream and uses an image analysis engine to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine and analyzes what the user is saying. The results of these analyses are combined to identify the problem the user is facing.
[0502] 4. Solution Generation
[0503] The server searches the database for the best solution for the identified problem, generates a video tutorial for solving the problem, creates an explanatory video, and generates a text description of the solution and a video link.
[0504] 5. Providing solutions
[0505] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and if the user clicks on the video, plays the explanatory video.
[0506] Specific examples
[0507] Example 1: If you cannot log in
[0508] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[0509] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[0510] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[0511] Example 2: When an error message appears
[0512] The user uses the camera to capture the situation where an error message is displayed and speaks. The device then sends this to the server.
[0513] The server extracts the specific error message from the video data and analyzes the user's utterances from the audio data. From these results, the server determines that the "specific error message" is the cause.
[0514] The server searches the database for a solution to this error message, generates an explanatory video showing the specific solution method, and sends the generated solution explanation and video link to the user, who then displays it on the device.
[0515] As described above, the present invention provides a system for quickly and accurately identifying problems faced by users and providing appropriate solutions, thereby significantly improving user satisfaction.
[0516] The processing flow will be explained below.
[0517] Step 1:
[0518] User: Launch the support app and click the "Start Video Call" button.
[0519] Step 2:
[0520] Device: Asks the user for permission to use the camera and microphone. If permission is granted, turns on the camera and microphone and generates the video call stream.
[0521] Step 3:
[0522] Terminal: Sends the video call stream to the server. Video and audio data are encoded in real time and continuously transmitted to the server.
[0523] Step 4:
[0524] Server: Receives connection requests from users' devices and creates new video call sessions. It buffers the received video and audio streams.
[0525] Step 5:
[0526] Server: Extracts each frame from the received video stream and uses an image analysis engine to identify GUI states and error messages.
[0527] Step 6:
[0528] Server: Converts the audio stream into text using a speech analysis engine. Analyzes what the user says and extracts keywords.
[0529] Step 7:
[0530] Server: Combines the results of video and audio analysis, applies rule-based filters, and identifies the problem the user is facing.
[0531] Step 8:
[0532] Server: Searches for a suitable solution from a database based on the identified problem, and generates an instructional video demonstrating the solution if necessary.
[0533] Step 9:
[0534] Server: Generates a text description of the solution and a video link and sends them to the user's device.
[0535] Step 10:
[0536] On the device: The received text description and video link are displayed to the user. If the user clicks on the video link, the video player is launched and the explanatory video is played.
[0537] Step 11:
[0538] User: Check the solution displayed on the device and take the necessary action. If the solution is insufficient, ask for additional support via video call again.
[0539] Example 1
[0540] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0541] Conventional user support systems require users to describe their problems in detail, a process that takes time and effort. Furthermore, identifying problems and providing solutions is often done manually, which is often inefficient. These issues result in lower user satisfaction and increased support costs.
[0542] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0543] In this invention, the server includes means for analyzing the status of the graphical user interface and error messages from the received video data, means for analyzing the content of user comments from the received audio data, and means for identifying the user's problem by integrating the results of the analysis of the video data and the audio data, thereby making it possible to efficiently and quickly identify the problem and provide a solution without the user having to explain the problem in detail.
[0544] The "means for initiating a video call" refers to the operation performed by a user to initiate a video call using an application on a terminal, and the combination of software and hardware that realizes this operation.
[0545] The "means for capturing and transmitting video and audio" is a mechanism for capturing video and audio data in real time using the camera and microphone of the terminal, encoding this data, and transmitting it to the server.
[0546] The "means for analyzing the received video and audio data to identify the problem the user is facing" is a software process in which the server processes the received video and audio data using an analysis engine to identify the problem or issue the user is facing.
[0547] The "means for generating a solution" is a combination of software and hardware that allows the server to search for an appropriate solution from the database and generate a solution to provide to the user based on the searched solution.
[0548] The "means for transmitting the generated solution description and video link to the user's terminal" is a communication means for transmitting the server-generated text description of the solution and the associated video link to the user's terminal via a network.
[0549] The "means for displaying the received solution description and video link to the user" is a mechanism including a user interface for displaying the text description and video link of the solution received by the terminal from the server in a format that is easy for the user to view.
[0550] "Means for analyzing the status of the graphical user interface and error messages" refers to a software process in which the server analyzes video data to extract information such as on-screen buttons, menus, and error messages, and then uses this information to identify the user's operational status and problems.
[0551] "Means for analyzing the user's speech" refers to a software process by which the server converts the speech data into text using speech recognition technology and understands from the text what issue the user is reporting.
[0552] The "means for searching for an appropriate solution from the database" is a mechanism by which the server searches for and obtains the most appropriate solution for a specified problem from among the solutions stored in the database.
[0553] The "means for generating video tutorials" refers to video generation technology and software that allows the server to create videos that include specific steps to solve a user's problem.
[0554] The "means for transmitting to the user's terminal" refers to a communication means for transmitting the video tutorial and solution information generated by the server to the user's terminal via a network.
[0555] The "means for playing the video tutorial to the user" refers to a mechanism including a media player that allows the user to play and watch the video tutorial received by the user's terminal.
[0556] MODE FOR CARRYING OUT THE INVENTION
[0557] The present invention is a system for identifying problems users encounter when using a service and providing solutions to those problems. This system significantly improves support efficiency because it does not require users to describe the problem in detail.
[0558] Starting a video call
[0559] When a user launches a supporting app and clicks the "Start Video Call" button, the device requests permission to use the camera and microphone. If permission is granted, the device activates the camera and microphone, generates a video call stream, and sends it to the server. This stream uses technologies such as WebRTC. The server receives the connection request and creates a new video call session.
[0560] Video and audio capture
[0561] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server. The encoding is done using H.264 and AAC encoders. The server stores the received video and audio streams in a buffer and prepares them for analysis.
[0562] Data analysis
[0563] The server extracts each frame from the received video stream and uses an image analysis engine (e.g., OpenCV or TensorFlow) to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine (e.g., Google Speech-to-Text or IBM Watson Speech to Text) and analyzes what the user is saying. The results of this analysis are combined to identify the problem the user is facing.
[0564] Solution Generation
[0565] The server searches for the best solution for the identified problem from a database (e.g., MySQL or MongoDB). It then generates a video tutorial to solve the problem. Adobe After Effects and FFmpeg are used to generate the video. It then generates a text description of the solution and a video link.
[0566] Providing solutions
[0567] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user. When the user clicks the video link, the device plays the explanatory video.
[0568] Specific examples
[0569] Example 1: If you cannot log in
[0570] The user launches the support app and says, "I can't log in." The device uses its camera and microphone to send video and audio data to the server. The server identifies the blank fields on the login screen from the video data and analyzes the audio data. This allows the server to identify the problem: "I can't log in because the password field is blank." The server searches its database for a solution to the correct login method and generates a video tutorial showing the login procedure along with the message, "An entry is required in the password field." This is sent to the user's device, which then displays an explanation of the solution and a video link.
[0571] Example 2: When an error message appears
[0572] The user films the situation in which "an error message is displayed" and speaks about it. The device sends this to the server. The server extracts the specific error message from the video data and analyzes the user's statement from the audio data. As a result, the server determines that the "specific error message" is the cause. The server searches a database for a solution to this error message and generates an explanatory video showing a specific method of dealing with the problem. The generated solution explanation and video link are sent to the user, and the device displays them to the user.
[0573] Example prompts for generative AI models
[0574] Below is an example of a prompt sentence to input to the generative AI model.
[0575] 1. "A user initiates a video call. Video and audio are captured and sent to the server."
[0576] 2. "The server analyzes the video data received and checks for error messages and the status of the GUI."
[0577] 3. "Based on the identified problem, we generate optimal solutions and tutorial videos."
[0578] As described above, this system is able to quickly and accurately identify user problems and provide appropriate solutions, which is expected to improve user satisfaction and significantly improve support efficiency.
[0579] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0580] Step 1:
[0581] Starting a video call
[0582] The user launches the support app and clicks the "Start Video Call" button.
[0583] Input: User action.
[0584] The device will display a popup requesting permission to use the camera and microphone, and if the user allows it, the camera and microphone will be activated.
[0585] Input: User permission.
[0586] The device generates a video call stream and sends it to the server using WebRTC technology. The device sets the stream's protocol information and starts the stream.
[0587] Output: Video call stream.
[0588] The server receives the connection request and creates a new video call session. The server generates a session ID and sends a confirmation message to the device.
[0589] Output: A new video call session.
[0590] Step 2:
[0591] Video and audio capture
[0592] The device uses a camera to capture video data and a microphone to capture audio data.
[0593] Input: Video and audio data.
[0594] The captured video and audio data is encoded in real time using H.264 and AAC encoders.
[0595] The terminal transmits the encoded video and audio data to the server.
[0596] Output: Encoded video and audio data.
[0597] The server stores the received data in a buffer and prepares it for analysis. The server manages the buffer memory and checks the data integrity.
[0598] Output: The data stored in the buffer.
[0599] Step 3:
[0600] Data analysis
[0601] The server extracts each frame from the received video stream using a video analysis engine (e.g., OpenCV).
[0602] Input: Buffered video data.
[0603] The server uses an image analysis engine to identify GUI states (e.g., buttons, input fields, error messages).
[0604] Input: Video data for each frame.
[0605] Output: Data of the identified GUI element.
[0606] The server converts the audio stream into text using a speech analysis engine (e.g., Google Speech-to-Text), which is then analyzed by a natural language processing engine.
[0607] Input: Buffered audio data.
[0608] Output: The audio data converted to text.
[0609] The server combines the results of the video and audio analysis to identify the problem the user is facing.
[0610] Input: Parsed GUI data and text data.
[0611] Output: Identified issues.
[0612] Step 4:
[0613] Solution Generation
[0614] The server searches for a solution to the identified problem from a database that stores information about existing solutions.
[0615] Input: Identified problem.
[0616] Output: The solution found.
[0617] The server uses a video generation engine (e.g. Adobe After Effects) to generate a video tutorial for solving the problem. The video is automatically generated based on the steps you specify.
[0618] Input: The solution found.
[0619] Output: The generated video tutorial.
[0620] The server generates a message containing a text description of the solution and a video link.
[0621] Input: Retrieved solutions and generated video tutorials.
[0622] Output: Solution description and video link.
[0623] Step 5:
[0624] Providing solutions
[0625] The server generates a solution explanation and sends a video link to the user's device.
[0626] Input: Solution description and video link.
[0627] Output: Data sent to the user's terminal.
[0628] The device displays the received solution explanation and video link to the user.
[0629] Input: Data received from the server.
[0630] Output: The solution description and video link that is displayed to the user.
[0631] When the user clicks on the video link, the device will play the instructional video.
[0632] Input: Click action on the video link.
[0633] Output: The instruction video that will be played.
[0634] (Application example 1)
[0635] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0636] When a machine or robot operating in a factory begins to operate abnormally, it is necessary to quickly and accurately identify the problem and provide an appropriate solution. However, currently, it takes a lot of time for operators to identify the problem and find a solution. In addition, finding an appropriate solution requires specialized knowledge, which leads to a decrease in production efficiency. Therefore, there is a need to introduce a system that can detect abnormalities in real time, efficiently identify the problem, and provide a solution.
[0637] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0638] In this invention, the server includes a means for detecting abnormal operation of the machine, a means for transmitting data captured by the machine to the server, and a means for identifying an optimal solution for the machine and instructing the machine on how to deal with the problem. This makes it possible to quickly detect abnormal operation, efficiently identify the problem, and provide an appropriate solution.
[0639] "Means for a user to initiate a video call" refers to software or hardware functionality that allows a user to initiate a video call session to receive remote support.
[0640] "Means for the terminal to capture and transmit the user's video and audio" refers to a function that enables the terminal to use a camera and microphone to capture the user's video and audio and transmit them to a server.
[0641] "Means for analyzing the video and audio data received by the server to identify the problem the user is facing" is a function for analyzing the video and audio data received by the server and identifying the problem the user is facing.
[0642] The "means by which the server generates a solution to the identified problem" is a function by which the server generates an appropriate solution to the problem after identifying the problem.
[0643] "Means for transmitting the explanation of the solution generated by the server and a video link to the user's terminal" is a function for transmitting the explanation of the solution generated by the server and a related video link to the user's terminal.
[0644] The "means for displaying to the user the explanation of the solution and the video link received by the terminal" is a function for displaying to the user the explanation of the solution and the video link received by the user's terminal.
[0645] "Means for machines to detect abnormal operation" refers to software or hardware functions that allow machines operating in a factory to monitor their operating status and detect any abnormalities that may occur.
[0646] "Means for transmitting data captured by the machine to the server" is a function for transmitting video and audio data captured by the machine to the server when an abnormality is detected.
[0647] "Means for identifying the optimal solution for the machine and instructing the machine on how to respond" is a function that enables the server to analyze a problem, identify an appropriate solution, and instruct the machine on how to respond.
[0648] MODE FOR CARRYING OUT THE INVENTION
[0649] A system for implementing this invention includes various means for quickly and accurately identifying a user's problem and providing an appropriate solution. Specifically, a user initiates a video call, and the video and audio data captured by the terminal is sent to a server, which analyzes the data and identifies the problem. The server then generates a solution and provides it to the user. In a factory, if a machine detects an abnormal operation, it notifies the server, which then analyzes the situation and provides an optimal solution.
[0650] Hardware and Software
[0651] Factory robots: equipped with cameras, microphones, and communication modules, including sensors to detect abnormal behavior.
[0652] Server: A video analysis engine (e.g., OpenCV), an audio analysis engine (e.g., Google Speech-to-Text API), a database (e.g., MySQL), and a solution generation module (Python script).
[0653] Device: Includes the device used by the user (smartphone, tablet, PC, etc.).
[0654] Data processing and calculation
[0655] The server first receives video and audio data sent from terminals and factory robots. The received video data is analyzed using an image analysis engine (OpenCV) to identify GUI status and error messages. Meanwhile, the audio data is converted into text using a speech analysis engine (Google Speech-to-Text API), and the user's speech is analyzed.
[0656] The analyzed data is integrated and the server uses this information to identify the problems the user is facing. For the identified problems, the server searches the database (MySQL) for the best solution and generates a text explanation and a video tutorial as the solution.
[0657] The generated solution is sent from the server to the user's device, where it is displayed to the user, along with notifications and explanatory video links sent from the device.
[0658] Specific examples
[0659] 1. Resolving login issues:
[0660] The user says, "I can't log in."
[0661] The device uses a camera and microphone to capture video and audio and transmits them to a server.
[0662] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[0663] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays the solution explanation and video link to the user.
[0664] Example prompt sentence:
[0665] Identify issues from video call data and provide solutions.
[0666] Audio data: "I can't log in"
[0667] Video data: Screenshot of the login screen
[0668] 2. Machine Anomaly Detection and Resolution:
[0669] A factory machine detects abnormal operation.
[0670] The machine uses a camera and microphone to capture video and audio and transmits it to a server.
[0671] The server identifies specific error messages and abnormal behavior patterns from the video data and also analyzes information from the audio data. Based on the results, the server identifies the cause of the abnormality and searches a database for an appropriate solution.
[0672] The server instructs the machine on the generated solutions and workarounds, allowing it to make the necessary corrections.
[0673] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0674] Step 1:
[0675] The user launches the supporting app on the device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone. The input obtained here is the user's video and audio, which are generated as a video call stream and sent to the server.
[0676] Step 2:
[0677] The server buffers the video call stream received from the device. The received video data is divided into frames, and the audio data is encoded. The input to this analysis preparation stage is the video and audio streams, and the output is data ready for analysis.
[0678] Step 3:
[0679] The server uses an image analysis engine (OpenCV) to analyze each frame of the received video stream. Specifically, it performs processing to identify the GUI status and error messages. The input is the received video data, and the output is information about the GUI status and error messages.
[0680] Step 4:
[0681] The server converts the received audio stream into text using a speech analysis engine (Google Speech-to-Text API). This converts the user's input voice data into text data and analyzes the user's speech. The input is voice data, and the output is text data.
[0682] Step 5:
[0683] The server integrates the results of the video data analysis obtained in step 3 and the results of the audio data analysis obtained in step 4. This identifies the problem the user is facing. The inputs are the video analysis results and the audio analysis results, and the output is information about the identified problem.
[0684] Step 6:
[0685] The server searches for a suitable solution from a database (MySQL) based on the identified problem. In this step, the input is information about the identified problem and the output is the solution retrieved from the database.
[0686] Step 7:
[0687] The server generates a video tutorial for solving the problem based on the solution to the identified problem. The generated solution includes a text explanation and a video link. The input is the searched solution, and the output is the video tutorial and its link.
[0688] Step 8:
[0689] The server sends the generated solution explanation and video link to the user's device. The input is the video tutorial and link, and the output is the data sent to the user's device.
[0690] Step 9:
[0691] The device displays the received solution explanation and video link to the user. Specifically, when the user clicks on the video link, the explanatory video is played. The input is the received data, and the output is the solution explanation and video link displayed to the user.
[0692] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0693] The present invention is a system that identifies problems users encounter when using services, taking into account the user's emotional state, and provides solutions. Because the user does not need to explain the problem in detail, support efficiency can be significantly improved. A specific embodiment of this system is described below.
[0694] Overall flow
[0695] 1. Start a video call
[0696] A user launches a supporting app on their device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone, generates a video call stream, and sends it to the server. The server receives the connection request and creates a new video call session.
[0697] 2. Video and audio capture
[0698] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server, which stores the received video and audio streams in a buffer and prepares them for analysis.
[0699] 3. Data Analysis
[0700] The server extracts each frame from the received video stream and uses an image analysis engine to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine and analyzes what the user is saying. The results of these analyses are combined to identify the problem the user is facing.
[0701] 4. Emotion analysis
[0702] The server uses an emotion engine that recognizes the user's emotions from the video and audio data, for example, facial expressions and tone of voice to determine whether the user is feeling impatient, angry, sad, etc.
[0703] 5. Solution Generation
[0704] The server searches for the best solution from the database based on the identified problem and the user's emotional state. If necessary, it adjusts the priority and presentation of the solutions taking the emotional state into account. It also generates an appropriate video tutorial for solving the problem and creates an explanatory video. It also generates a text description and video link for the solution.
[0705] 6. Providing solutions
[0706] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and if the user clicks on the video, the device launches a video player and plays the explanatory video.
[0707] Specific examples
[0708] Example 1: If you cannot log in
[0709] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[0710] Emotion analysis: The server determines from video and audio data that the user is feeling anxious, and determines that a solution needs to be offered quickly.
[0711] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[0712] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[0713] Example 2: When an error message appears
[0714] The user uses the camera to capture the situation where an error message is displayed and speaks. The device then sends this to the server.
[0715] Sentiment analysis: The server determines from video and audio data that the user is angry, and determines that a solution should be provided in a calm and easy-to-understand format.
[0716] The server extracts the specific error message from the video data and analyzes the user's utterances from the audio data. From these results, the server determines that the "specific error message" is the cause.
[0717] The server searches the database for a solution to this error message, generates an explanatory video showing the specific solution method, and sends the generated solution explanation and video link to the user, who then displays it on the device.
[0718] As described above, the present invention provides a system that quickly and accurately identifies the problems a user is facing, taking into account the user's emotional state, and provides appropriate solutions, thereby significantly improving user satisfaction.
[0719] The processing flow will be explained below.
[0720] Step 1:
[0721] User: Launch the support app and click the "Start Video Call" button.
[0722] Step 2:
[0723] Device: Asks the user for permission to use the camera and microphone. If permission is granted, turns on the camera and microphone.
[0724] Step 3:
[0725] Terminal: Generates the video call stream, encodes the video and audio data in real time, and sends it to the server.
[0726] Step 4:
[0727] Server: Receives connection requests from users' devices and creates new video call sessions. It buffers the received video and audio streams.
[0728] Step 5:
[0729] Server: Extracts each frame from the video stream and uses an image analysis engine to identify GUI states and error messages.
[0730] Step 6:
[0731] Server: The audio stream is converted into text using a speech analysis engine, and the content of what the user is saying is analyzed.
[0732] Step 7:
[0733] Server: Integrates the results of video and audio analysis and applies rule-based filters to identify the issues the user is facing.
[0734] Step 8:
[0735] Server: Uses an emotion engine to recognize the user's emotions from video and audio data. Emotions are analyzed from the user's facial expressions and tone of voice.
[0736] Step 9:
[0737] Server: Searches for an appropriate solution from a database based on the identified problem and the user's emotional state.
[0738] Step 10:
[0739] Server: Tailors solutions to user emotions and prioritizes them as needed. Generates problem-solving video tutorials and creates instructional videos.
[0740] Step 11:
[0741] Server: Generates a text description of the solution and a video link and sends them to the user's device.
[0742] Step 12:
[0743] On your device: Display the received text description and video link to the user.
[0744] Step 13:
[0745] User: Clicks on the video link and watches a video that explains how to solve the problem.
[0746] Step 14:
[0747] Device: Plays instructional videos and helps users solve problems by following instructions.
[0748] The above steps create a system that can quickly and accurately resolve problems users encounter while using the service. The introduction of an emotion engine makes it possible to provide optimal support based on the user's emotional state.
[0749] Example 2
[0750] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0751] In today's world, users face a wide variety of problems when using services. Support for resolving these problems typically requires users to explain the problem in detail. However, if users are unable to clearly express their problems, solving the problem increases the time and effort required. Furthermore, presenting solutions without considering the user's emotional state can detract from the user experience. To address these challenges, a system is needed that can quickly and accurately identify problems and provide optimal solutions, taking the user's emotions into account.
[0752] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0753] In this invention, the server includes means for analyzing the user's emotions from the received video data and audio data, means for generating an appropriate solution based on the identified problem and the user's emotional state, and means for transmitting an explanation of the generated solution and a video link to the user's terminal. This makes it possible to quickly and accurately identify the problem and provide an optimal solution while taking into account the user's emotions, without the user having to clearly explain the problem.
[0754] "User" means a person who uses the System to request support for the Service.
[0755] "Video call" refers to a real-time means of video and audio communication between a user and a support system.
[0756] A "terminal" is a device that a user uses to interact with the support system, such as a smartphone or computer.
[0757] A "camera" is a device for capturing images that is built into or attached to a terminal.
[0758] A "microphone" is a device for capturing sound that is built into a terminal or attached externally.
[0759] "Server" is a central processing system that analyzes data received through the video call and provides appropriate solutions to the user.
[0760] "Video data" refers to image information captured by a camera.
[0761] "Audio Data" means audio information captured by a microphone.
[0762] "Analysis" is the act of extracting and processing information based on video and audio data.
[0763] "Emotion analysis" is the process of identifying a user's emotional state based on information such as facial expressions and tone of voice.
[0764] A "solution" is a specific procedure or method provided to solve a problem a user is facing.
[0765] A "video link" is a hyperlink that indicates where to watch an explanatory video, etc.
[0766] "Display" is the act of visualizing information such as text or video links on a device screen.
[0767] A "video player" is software or an application that plays an explanatory video when you click on a video link.
[0768] MODE FOR CARRYING OUT THE INVENTION
[0769] The present invention is a system for identifying problems users encounter when using services, taking into account the user's emotional state, and providing solutions, allowing users to receive prompt and accurate support without having to explain their problems in detail.
[0770] Hardware and software used
[0771] Hardware
[0772] Device: The smartphone or computer used by the user.
[0773] Camera: A built-in or external video capture device.
[0774] Microphone: Built-in or external audio capture device.
[0775] Server: A central processing system that processes and analyzes data.
[0776] software
[0777] Supporting App: An application that allows users to initiate video calls.
[0778] Image analysis engine: For example, OpenCV is used to analyze video data.
[0779] Speech analysis engine: Converts voice data into text using, for example, the Google Speech-to-Text API.
[0780] Sentiment analysis engine: Analyze the user's emotional state using, for example, the Affectiva SDK.
[0781] System action
[0782] Starting a video call
[0783] 1. The user launches the supporting app on their device and clicks the "Start Video Call" button.
[0784] 2. The device requests permission to use the camera and microphone, and if the user grants permission, it activates the camera and microphone, generates a video call stream, and sends it to the server.
[0785] 3. The server receives the connection request and creates a new video call session.
[0786] Video and audio capture
[0787] 1. The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server.
[0788] 2. The server buffers the received video and audio streams and prepares them for analysis.
[0789] Data analysis
[0790] 1. The server extracts each frame from the received video stream and uses an image analysis engine to identify the status of the graphical user interface and error messages.
[0791] 2. At the same time, the audio stream is converted into text by a speech analysis engine and the content of what the user is saying is analyzed.
[0792] 3. Synthesize the analysis results to identify the problems users are facing.
[0793] Emotion analysis
[0794] 1. The server uses an emotion analysis engine to recognize the user's emotions from video and audio data.
[0795] 2. For example, by looking at facial expressions and tone of voice, it can identify whether the user is feeling anxious, angry, sad, or other emotions.
[0796] Solution Generation
[0797] 1. The server searches the database for the best solution based on the identified problem and the user's emotional state.
[0798] 2. Adjust the priorities and presentation of solutions to take into account emotional states.
[0799] 3. Generate appropriate video tutorials and create instructional videos to solve problems.
[0800] 4. Create a text description and video link of the generated solution.
[0801] Providing solutions
[0802] 1. The server sends the generated solution explanation and video link to the user's device.
[0803] 2. The device displays the received text description and video link to the user.
[0804] 3. When the user clicks on the video link, the device launches a video player and plays the explanatory video.
[0805] Specific examples
[0806] Example 1: If you cannot log in
[0807] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[0808] Emotion analysis: The server determines from video and audio data that the user is feeling anxious, and determines that a solution needs to be offered quickly.
[0809] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data, thereby identifying the problem of "I can't log in because the password field is blank."
[0810] The server searches the database for a solution to the correct login method, generates a message saying "Password field must be filled in" and a video tutorial showing the login procedure, and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[0811] Prompt Sentence Examples
[0812] "Users report issues via video call, you analyze them including sentiment analysis, and provide solutions with video tutorials. Users say 'I can't log in.'"
[0813] The above is an example of the present invention. This system enables the user to quickly and accurately identify problems and provide optimal solutions, taking into account emotions, without the user having to clearly explain the problem.
[0814] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0815] Specific explanation of processing steps
[0816] Step 1:
[0817] Starting a video call
[0818] Input: The user launches a supporting app and clicks the "Start Video Call" button.
[0819] How it works: A user launches a supporting app and clicks the button for a video call.
[0820] Output: The device displays a dialog to allow use of the camera and microphone.
[0821] Example of how it works:
[0822] When the user clicks the "Start Video Call" button, a dialog box will appear requesting permission to use the device's camera and microphone. If the user clicks "Allow," the camera and microphone will be activated.
[0823] Step 2:
[0824] Video and audio capture
[0825] Input: Camera and microphone activated.
[0826] How it works: The device captures video with its camera and audio with its microphone.
[0827] Output: Encoded video and audio data is generated and sent to the server.
[0828] Example of how it works:
[0829] The user's face and voice are captured in real time by a camera and microphone, encoded, and then streamed to a server.
[0830] Step 3:
[0831] Receiving and buffering data
[0832] Input: Video and audio data sent to the server.
[0833] Operation: The server stores the received video and audio data in a temporary buffer and prepares it for analysis.
[0834] Output: The data stored in the buffer is passed to the analysis engine.
[0835] Example of how it works:
[0836] When the server receives the video and audio streaming data, it is temporarily stored in a buffer, through which the data is passed to the next analysis step.
[0837] Step 4:
[0838] Data analysis
[0839] Input: Buffered video and audio data.
[0840] How it works: The server uses an image analysis engine (e.g., OpenCV) to analyze each frame of video data and recognize graphical user interfaces and error messages, while simultaneously using a speech analysis engine (e.g., Google Speech-to-Text API) to convert the audio data into text and analyze what the user is saying.
[0841] Output: Text data of the identified problem and what the user said.
[0842] Example of how it works:
[0843] A specific error message is read from the video data, and the statement "I can't log in" is converted into text from the audio data.
[0844] Step 5:
[0845] Emotion analysis
[0846] Input: Buffered video and audio data, and the identified issue.
[0847] How it works: The server uses an emotion analysis engine (e.g., Affectiva SDK) to analyze the user's emotional state, for example, by identifying the user's emotions (annoyance, anger, sadness, etc.) from facial expressions and tone of voice.
[0848] Output: Analysis data of the user's emotional state.
[0849] Example of how it works:
[0850] A video of the user's impatient face after seeing the identified error message is analyzed to identify the emotion of impatience.
[0851] Step 6:
[0852] Solution Generation
[0853] Input: Analysis data of identified problems and emotional states.
[0854] How it works: The server searches for the best solution from a database based on the identified problem and the user's emotional state, adjusts the priority and presentation of solutions as needed, and generates a video tutorial for solving the problem and creates an instructional video.
[0855] Output: A text description of the solution and a video link.
[0856] Example of how it works:
[0857] The server generates a "Password field is required" message and a video tutorial showing login steps.
[0858] Step 7:
[0859] Providing solutions
[0860] Input: A text description of the solution and a video link.
[0861] Operation: The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and when the user clicks the video link, it launches a video player and plays the explanatory video.
[0862] Output: The solution shown to the user and a video link.
[0863] Example of how it works:
[0864] The device screen displays the message "Password field requires input" along with a video link. When the user clicks the link, a video player launches and a video with the solution is played.
[0865] (Application example 2)
[0866] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0867] When users encounter security-related problems, conventional support systems often delay problem identification and solution provision, making it difficult to provide a prompt and appropriate response to the user. Furthermore, users who are in an emotional state, such as frustration, anger, or sadness, may find it difficult to understand the support content. This tends to reduce user satisfaction and prolong the problem. Therefore, there is a need for a system that not only speeds up problem identification but also provides solutions that take into account the user's emotional state.
[0868] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing received video and audio data to identify a problem faced by the user, means for analyzing the user's emotional state from the received video and audio data, and means for generating a solution based on the identified problem and the user's emotional state. This makes it possible, when a user faces a security-related problem, to quickly identify the problem and provide an effective solution that takes the user's emotional state into consideration.
[0869] "Video calling" is a function that allows users to communicate remotely in real time using audio and video.
[0870] A "terminal" is a device that a user uses to initiate a video call and capture video and audio. Examples include a smartphone or tablet.
[0871] "Server" is a central processing unit that analyzes video and audio data received from users, identifies problems, and provides solutions.
[0872] "Video data" refers to video information captured by a camera, including the user's face and the state of the screen interface.
[0873] "Audio data" is acoustic information captured by a microphone, including the user's voice and what is being said.
[0874] "Analysis" is the process by which the server extracts useful information from the data it receives and understands its meaning.
[0875] "Emotional state" refers to the emotions that a user expresses through video and audio data, such as impatience, anger, sadness, etc.
[0876] A "solution" is a method, procedure, or supporting information provided by the server to solve an identified problem.
[0877] "Description and Video Link" is a textual summary of the solution and access to a tutorial video related to that solution.
[0878] "Database" means an information management system that accumulates and stores solutions and related information in a searchable format.
[0879] "Video tutorials" are video contents that visually explain to users how to solve problems.
[0880] Overall flow
[0881] The invention includes a system for quickly and effectively responding when a user experiences a security-related problem, which takes into account the user's emotional state to identify the problem and provide a solution.
[0882] Hardware and Software
[0883] Hardware used:
[0884] Smartphone or tablet (e.g. iPhone, Android device)
[0885] Server (high-performance server for sensing and analysis, such as Amazon Web Services)
[0886] Software used:
[0887] Video calling libraries (e.g. WebRTC)
[0888] Sentiment analysis engine (e.g. Microsoft Azure Emotion API)
[0889] Speech analysis engine (e.g. Google Cloud Speech-to-Text)
[0890] A database (e.g., MySQL or PostgreSQL)
[0891] Video platform (e.g. YouTube API)
[0892] System Operation
[0893] 1. Start a video call
[0894] A user launches a specific supporting application on a smartphone or tablet. This application uses WebRTC to provide video calling functionality. When the user clicks the "Start Video Call" button, the camera and microphone are activated and video and audio capture begins.
[0895] 2. Capture of video and audio data
[0896] Video data captured by the smartphone camera and audio data captured by the microphone are encoded in real time and sent to a server, which stores the data in a buffer and prepares it for analysis.
[0897] 3. Data Analysis
[0898] The server divides the received video stream into frames and identifies the face and screen interface state. Libraries such as OpenCV are used for image analysis. Additionally, audio data is converted into text using Google Cloud Speech-to-Text, allowing the user's speech to be analyzed.
[0899] 4. Emotion analysis
[0900] The server uses the received video and audio data to determine the user's emotional state. It uses the Microsoft Azure Emotion API to analyze emotions from facial expressions and complements emotions from the user's tone of voice.
[0901] 5. Solution Generation
[0902] The server searches for an appropriate solution from a database based on the identified problem and the user's emotional state, and then generates a problem-solving explanation text and related tutorial video using the YouTube API.
[0903] 6. Providing solutions
[0904] The server sends the generated solution explanation and video link to the user's device, which displays it to the user. When the user clicks on the video link, a video player is launched and the explanatory video is played.
[0905] Specific examples
[0906] Sentiment analysis prompt example:
[0907] Identify the user's emotions from their facial expressions and tone of voice, specifically whether they are anger, sadness, confusion, or impatience.
[0908] Solution search prompt example:
[0909] Find a solution to the security issue based on the problem the user is facing and their emotional state. Describe your recommended solution and steps, and provide links to relevant video tutorials.
[0910] Thus, an embodiment of the invention provides a system that quickly and accurately solves security problems while taking into account the emotional state of the user.
[0911] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0912] Step 1:
[0913] A user launches a supporting application and clicks the "Start Video Call" button. This receives user input and activates the camera and microphone as output. This prepares the application for capturing video and audio. Specifically, the application uses the WebRTC library to request permission to use the camera and microphone.
[0914] Step 2:
[0915] The device begins capturing video data with the camera and audio data with the microphone. The captured video and audio input data is encoded and sent to the server in real time. The specific operation is a process that uses WebRTC to generate video and audio streams and send the data to the server. The output is the video and audio data sent to the server.
[0916] Step 3:
[0917] The server splits the video stream into frames and uses an image analysis engine (e.g., OpenCV) to determine the GUI state and error messages. This is the process of analyzing the captured video data (input) and extracting specific screen elements and messages (output). The specific operation is to apply image processing algorithms and identify important elements.
[0918] Step 4:
[0919] The server converts the audio stream into text using Google Cloud Speech-to-Text. The input is the transmitted audio data, and the output is the analyzed text data. The specific operation is to send the audio data to the speech analysis engine and obtain the resulting text.
[0920] Step 5:
[0921] The server integrates the analysis results of the video and audio data to identify the problem the user is facing. This is the process of combining the outputs from the image analysis engine and the audio analysis engine to identify the problem. Specifically, it centrally manages the analysis results and applies an algorithm to identify the cause of the problem. The input is the image analysis results and the audio analysis results, and the output is the identified problem.
[0922] Step 6:
[0923] The server performs emotion analysis from the video and audio data. It uses the Microsoft Azure Emotion API to identify the user's emotional state. At this stage, the input is the user's video and audio data, and the output is the analyzed emotional state. The specific operation is to send the data to the emotion analysis engine and obtain the resulting emotional information.
[0924] Step 7:
[0925] The server searches for an appropriate solution from a database based on the identified problem and the user's emotional state. The input is the analysis results and the user's emotional state, and the output is an appropriate solution. The specific operation is to quickly extract a solution using a database search algorithm.
[0926] Step 8:
[0927] The server generates an explanatory video to solve the problem. It also generates links to related tutorial videos using the YouTube API. The input is the searched solution, and the output is a solution explanation including a video link. The specific operation is the process of linking with the YouTube API and extracting the appropriate video.
[0928] Step 9:
[0929] The server sends the generated solution explanation and video link to the user's device. The input is the solution explanation and video link, and the output is data transmission to the user's device. The specific operation is to send data to the user's device via network communication.
[0930] Step 10:
[0931] The device displays the received solution explanation and video link to the user. When the user clicks on the video link, a video player on the device is launched and the explanatory video is played. The input is the solution explanation and video link sent from the server, and the output is the display and video playback on the device. The specific operation is to display the solution on the user interface and launch the video player.
[0932] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0933] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0934] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0935] [Third embodiment]
[0936] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0937] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0938] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0939] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0940] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0941] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0942] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0943] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0944] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0945] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0946] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0947] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0948] The present invention is a system for identifying problems that users encounter when using a service and providing solutions to those problems. This system significantly improves support efficiency because it eliminates the need for users to explain the problem in detail. A specific embodiment of this system is described below.
[0949] Overall flow
[0950] 1. Start a video call
[0951] A user launches a supporting app on their device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone, generates a video call stream, and sends it to the server. The server receives the connection request and creates a new video call session.
[0952] 2. Video and audio capture
[0953] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server, which stores the received video and audio streams in a buffer and prepares them for analysis.
[0954] 3. Data Analysis
[0955] The server extracts each frame from the received video stream and uses an image analysis engine to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine and analyzes what the user is saying. The results of these analyses are combined to identify the problem the user is facing.
[0956] 4. Solution Generation
[0957] The server searches the database for the best solution for the identified problem, generates a video tutorial for solving the problem, creates an explanatory video, and generates a text description of the solution and a video link.
[0958] 5. Providing solutions
[0959] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and if the user clicks on the video, plays the explanatory video.
[0960] Specific examples
[0961] Example 1: If you cannot log in
[0962] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[0963] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[0964] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[0965] Example 2: When an error message appears
[0966] The user uses the camera to capture the situation where an error message is displayed and speaks. The device then sends this to the server.
[0967] The server extracts the specific error message from the video data and analyzes the user's utterances from the audio data. From these results, the server determines that the "specific error message" is the cause.
[0968] The server searches the database for a solution to this error message, generates an explanatory video showing the specific solution method, and sends the generated solution explanation and video link to the user, who then displays it on the device.
[0969] As described above, the present invention provides a system for quickly and accurately identifying problems faced by users and providing appropriate solutions, thereby significantly improving user satisfaction.
[0970] The processing flow will be explained below.
[0971] Step 1:
[0972] User: Launch the support app and click the "Start Video Call" button.
[0973] Step 2:
[0974] Device: Asks the user for permission to use the camera and microphone. If permission is granted, turns on the camera and microphone and generates the video call stream.
[0975] Step 3:
[0976] Terminal: Sends the video call stream to the server. Video and audio data are encoded in real time and continuously transmitted to the server.
[0977] Step 4:
[0978] Server: Receives connection requests from users' devices and creates new video call sessions. It buffers the received video and audio streams.
[0979] Step 5:
[0980] Server: Extracts each frame from the received video stream and uses an image analysis engine to identify GUI states and error messages.
[0981] Step 6:
[0982] Server: Converts the audio stream into text using a speech analysis engine. Analyzes what the user says and extracts keywords.
[0983] Step 7:
[0984] Server: Combines the results of video and audio analysis, applies rule-based filters, and identifies the problem the user is facing.
[0985] Step 8:
[0986] Server: Searches for a suitable solution from a database based on the identified problem, and generates an instructional video demonstrating the solution if necessary.
[0987] Step 9:
[0988] Server: Generates a text description of the solution and a video link and sends them to the user's device.
[0989] Step 10:
[0990] On the device: The received text description and video link are displayed to the user. If the user clicks on the video link, the video player is launched and the explanatory video is played.
[0991] Step 11:
[0992] User: Check the solution displayed on the device and take the necessary action. If the solution is insufficient, ask for additional support via video call again.
[0993] Example 1
[0994] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0995] Conventional user support systems require users to describe their problems in detail, a process that takes time and effort. Furthermore, identifying problems and providing solutions is often done manually, which is often inefficient. These issues result in lower user satisfaction and increased support costs.
[0996] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0997] In this invention, the server includes means for analyzing the status of the graphical user interface and error messages from the received video data, means for analyzing the content of user comments from the received audio data, and means for identifying the user's problem by integrating the results of the analysis of the video data and the audio data, thereby making it possible to efficiently and quickly identify the problem and provide a solution without the user having to explain the problem in detail.
[0998] The "means for initiating a video call" refers to the operation performed by a user to initiate a video call using an application on a terminal, and the combination of software and hardware that realizes this operation.
[0999] The "means for capturing and transmitting video and audio" is a mechanism for capturing video and audio data in real time using the camera and microphone of the terminal, encoding this data, and transmitting it to the server.
[1000] The "means for analyzing the received video and audio data to identify the problem the user is facing" is a software process in which the server processes the received video and audio data using an analysis engine to identify the problem or issue the user is facing.
[1001] The "means for generating a solution" is a combination of software and hardware that allows the server to search for an appropriate solution from the database and generate a solution to provide to the user based on the searched solution.
[1002] The "means for transmitting the generated solution description and video link to the user's terminal" is a communication means for transmitting the server-generated text description of the solution and the associated video link to the user's terminal via a network.
[1003] The "means for displaying the received solution description and video link to the user" is a mechanism including a user interface for displaying the text description and video link of the solution received by the terminal from the server in a format that is easy for the user to view.
[1004] "Means for analyzing the status of the graphical user interface and error messages" refers to a software process in which the server analyzes video data to extract information such as on-screen buttons, menus, and error messages, and then uses this information to identify the user's operational status and problems.
[1005] "Means for analyzing the user's speech" refers to a software process by which the server converts the speech data into text using speech recognition technology and understands from the text what issue the user is reporting.
[1006] The "means for searching for an appropriate solution from the database" is a mechanism by which the server searches for and obtains the most appropriate solution for a specified problem from among the solutions stored in the database.
[1007] The "means for generating video tutorials" refers to video generation technology and software that allows the server to create videos that include specific steps to solve a user's problem.
[1008] The "means for transmitting to the user's terminal" refers to a communication means for transmitting the video tutorial and solution information generated by the server to the user's terminal via a network.
[1009] The "means for playing the video tutorial to the user" refers to a mechanism including a media player that allows the user to play and watch the video tutorial received by the user's terminal.
[1010] MODE FOR CARRYING OUT THE INVENTION
[1011] The present invention is a system for identifying problems users encounter when using a service and providing solutions to those problems. This system significantly improves support efficiency because it does not require users to describe the problem in detail.
[1012] Starting a video call
[1013] When a user launches a supporting app and clicks the "Start Video Call" button, the device requests permission to use the camera and microphone. If permission is granted, the device activates the camera and microphone, generates a video call stream, and sends it to the server. This stream uses technologies such as WebRTC. The server receives the connection request and creates a new video call session.
[1014] Video and audio capture
[1015] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server. The encoding is done using H.264 and AAC encoders. The server stores the received video and audio streams in a buffer and prepares them for analysis.
[1016] Data analysis
[1017] The server extracts each frame from the received video stream and uses an image analysis engine (e.g., OpenCV or TensorFlow) to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine (e.g., Google Speech-to-Text or IBM Watson Speech to Text) and analyzes what the user is saying. The results of this analysis are combined to identify the problem the user is facing.
[1018] Solution Generation
[1019] The server searches for the best solution for the identified problem from a database (e.g., MySQL or MongoDB). It then generates a video tutorial to solve the problem. Adobe After Effects and FFmpeg are used to generate the video. It then generates a text description of the solution and a video link.
[1020] Providing solutions
[1021] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user. When the user clicks the video link, the device plays the explanatory video.
[1022] Specific examples
[1023] Example 1: If you cannot log in
[1024] The user launches the support app and says, "I can't log in." The device uses its camera and microphone to send video and audio data to the server. The server identifies the blank fields on the login screen from the video data and analyzes the audio data. This allows the server to identify the problem: "I can't log in because the password field is blank." The server searches its database for a solution to the correct login method and generates a video tutorial showing the login procedure along with the message, "An entry is required in the password field." This is sent to the user's device, which then displays an explanation of the solution and a video link.
[1025] Example 2: When an error message appears
[1026] The user films the situation in which "an error message is displayed" and speaks about it. The device sends this to the server. The server extracts the specific error message from the video data and analyzes the user's statement from the audio data. As a result, the server determines that the "specific error message" is the cause. The server searches a database for a solution to this error message and generates an explanatory video showing a specific method of dealing with the problem. The generated solution explanation and video link are sent to the user, and the device displays them to the user.
[1027] Example prompts for generative AI models
[1028] Below is an example of a prompt sentence to input to the generative AI model.
[1029] 1. "A user initiates a video call. Video and audio are captured and sent to the server."
[1030] 2. "The server analyzes the video data received and checks for error messages and the status of the GUI."
[1031] 3. "Based on the identified problem, we generate optimal solutions and tutorial videos."
[1032] As described above, this system is able to quickly and accurately identify user problems and provide appropriate solutions, which is expected to improve user satisfaction and significantly improve support efficiency.
[1033] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1034] Step 1:
[1035] Starting a video call
[1036] The user launches the support app and clicks the "Start Video Call" button.
[1037] Input: User action.
[1038] The device will display a popup requesting permission to use the camera and microphone, and if the user allows it, the camera and microphone will be activated.
[1039] Input: User permission.
[1040] The device generates a video call stream and sends it to the server using WebRTC technology. The device sets the stream's protocol information and starts the stream.
[1041] Output: Video call stream.
[1042] The server receives the connection request and creates a new video call session. The server generates a session ID and sends a confirmation message to the device.
[1043] Output: A new video call session.
[1044] Step 2:
[1045] Video and audio capture
[1046] The device uses a camera to capture video data and a microphone to capture audio data.
[1047] Input: Video and audio data.
[1048] The captured video and audio data is encoded in real time using H.264 and AAC encoders.
[1049] The terminal transmits the encoded video and audio data to the server.
[1050] Output: Encoded video and audio data.
[1051] The server stores the received data in a buffer and prepares it for analysis. The server manages the buffer memory and checks the data integrity.
[1052] Output: The data stored in the buffer.
[1053] Step 3:
[1054] Data analysis
[1055] The server extracts each frame from the received video stream using a video analysis engine (e.g., OpenCV).
[1056] Input: Buffered video data.
[1057] The server uses an image analysis engine to identify GUI states (e.g., buttons, input fields, error messages).
[1058] Input: Video data for each frame.
[1059] Output: Data of the identified GUI element.
[1060] The server converts the audio stream into text using a speech analysis engine (e.g., Google Speech-to-Text), which is then analyzed by a natural language processing engine.
[1061] Input: Buffered audio data.
[1062] Output: The audio data converted to text.
[1063] The server combines the results of the video and audio analysis to identify the problem the user is facing.
[1064] Input: Parsed GUI data and text data.
[1065] Output: Identified issues.
[1066] Step 4:
[1067] Solution Generation
[1068] The server searches for a solution to the identified problem from a database that stores information about existing solutions.
[1069] Input: Identified problem.
[1070] Output: The solution found.
[1071] The server uses a video generation engine (e.g. Adobe After Effects) to generate a video tutorial for solving the problem. The video is automatically generated based on the steps you specify.
[1072] Input: The solution found.
[1073] Output: The generated video tutorial.
[1074] The server generates a message containing a text description of the solution and a video link.
[1075] Input: Retrieved solutions and generated video tutorials.
[1076] Output: Solution description and video link.
[1077] Step 5:
[1078] Providing solutions
[1079] The server generates a solution explanation and sends a video link to the user's device.
[1080] Input: Solution description and video link.
[1081] Output: Data sent to the user's terminal.
[1082] The device displays the received solution explanation and video link to the user.
[1083] Input: Data received from the server.
[1084] Output: The solution description and video link that is displayed to the user.
[1085] When the user clicks on the video link, the device will play the instructional video.
[1086] Input: Click action on the video link.
[1087] Output: The instruction video that will be played.
[1088] (Application example 1)
[1089] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1090] When a machine or robot operating in a factory begins to operate abnormally, it is necessary to quickly and accurately identify the problem and provide an appropriate solution. However, currently, it takes a lot of time for operators to identify the problem and find a solution. In addition, finding an appropriate solution requires specialized knowledge, which leads to a decrease in production efficiency. Therefore, there is a need to introduce a system that can detect abnormalities in real time, efficiently identify the problem, and provide a solution.
[1091] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1092] In this invention, the server includes a means for detecting abnormal operation of the machine, a means for transmitting data captured by the machine to the server, and a means for identifying an optimal solution for the machine and instructing the machine on how to deal with the problem. This makes it possible to quickly detect abnormal operation, efficiently identify the problem, and provide an appropriate solution.
[1093] "Means for a user to initiate a video call" refers to software or hardware functionality that allows a user to initiate a video call session to receive remote support.
[1094] "Means for the terminal to capture and transmit the user's video and audio" refers to a function that enables the terminal to use a camera and microphone to capture the user's video and audio and transmit them to a server.
[1095] "Means for analyzing the video and audio data received by the server to identify the problem the user is facing" is a function for analyzing the video and audio data received by the server and identifying the problem the user is facing.
[1096] The "means by which the server generates a solution to the identified problem" is a function by which the server generates an appropriate solution to the problem after identifying the problem.
[1097] "Means for transmitting the explanation of the solution generated by the server and a video link to the user's terminal" is a function for transmitting the explanation of the solution generated by the server and a related video link to the user's terminal.
[1098] The "means for displaying to the user the explanation of the solution and the video link received by the terminal" is a function for displaying to the user the explanation of the solution and the video link received by the user's terminal.
[1099] "Means for machines to detect abnormal operation" refers to software or hardware functions that allow machines operating in a factory to monitor their operating status and detect any abnormalities that may occur.
[1100] "Means for transmitting data captured by the machine to the server" is a function for transmitting video and audio data captured by the machine to the server when an abnormality is detected.
[1101] "Means for identifying the optimal solution for the machine and instructing the machine on how to respond" is a function that enables the server to analyze a problem, identify an appropriate solution, and instruct the machine on how to respond.
[1102] MODE FOR CARRYING OUT THE INVENTION
[1103] A system for implementing this invention includes various means for quickly and accurately identifying a user's problem and providing an appropriate solution. Specifically, a user initiates a video call, and the video and audio data captured by the terminal is sent to a server, which analyzes the data and identifies the problem. The server then generates a solution and provides it to the user. In a factory, if a machine detects an abnormal operation, it notifies the server, which then analyzes the situation and provides an optimal solution.
[1104] Hardware and Software
[1105] Factory robots: equipped with cameras, microphones, and communication modules, including sensors to detect abnormal behavior.
[1106] Server: A video analysis engine (e.g., OpenCV), an audio analysis engine (e.g., Google Speech-to-Text API), a database (e.g., MySQL), and a solution generation module (Python script).
[1107] Device: Includes the device used by the user (smartphone, tablet, PC, etc.).
[1108] Data processing and calculation
[1109] The server first receives video and audio data sent from terminals and factory robots. The received video data is analyzed using an image analysis engine (OpenCV) to identify GUI status and error messages. Meanwhile, the audio data is converted into text using a speech analysis engine (Google Speech-to-Text API), and the user's speech is analyzed.
[1110] The analyzed data is integrated and the server uses this information to identify the problems the user is facing. For the identified problems, the server searches the database (MySQL) for the best solution and generates a text explanation and a video tutorial as the solution.
[1111] The generated solution is sent from the server to the user's device, where it is displayed to the user, along with notifications and explanatory video links sent from the device.
[1112] Specific examples
[1113] 1. Resolving login issues:
[1114] The user says, "I can't log in."
[1115] The device uses a camera and microphone to capture video and audio and transmits them to a server.
[1116] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[1117] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays the solution explanation and video link to the user.
[1118] Example prompt sentence:
[1119] Identify issues from video call data and provide solutions.
[1120] Audio data: "I can't log in"
[1121] Video data: Screenshot of the login screen
[1122] 2. Machine Anomaly Detection and Resolution:
[1123] A factory machine detects abnormal operation.
[1124] The machine uses a camera and microphone to capture video and audio and transmits it to a server.
[1125] The server identifies specific error messages and abnormal behavior patterns from the video data and also analyzes information from the audio data. Based on the results, the server identifies the cause of the abnormality and searches a database for an appropriate solution.
[1126] The server instructs the machine on the generated solutions and workarounds, allowing it to make the necessary corrections.
[1127] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1128] Step 1:
[1129] The user launches the supporting app on the device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone. The input obtained here is the user's video and audio, which are generated as a video call stream and sent to the server.
[1130] Step 2:
[1131] The server buffers the video call stream received from the device. The received video data is divided into frames, and the audio data is encoded. The input to this analysis preparation stage is the video and audio streams, and the output is data ready for analysis.
[1132] Step 3:
[1133] The server uses an image analysis engine (OpenCV) to analyze each frame of the received video stream. Specifically, it performs processing to identify the GUI status and error messages. The input is the received video data, and the output is information about the GUI status and error messages.
[1134] Step 4:
[1135] The server converts the received audio stream into text using a speech analysis engine (Google Speech-to-Text API). This converts the user's input voice data into text data and analyzes the user's speech. The input is voice data, and the output is text data.
[1136] Step 5:
[1137] The server integrates the results of the video data analysis obtained in step 3 and the results of the audio data analysis obtained in step 4. This identifies the problem the user is facing. The inputs are the video analysis results and the audio analysis results, and the output is information about the identified problem.
[1138] Step 6:
[1139] The server searches for a suitable solution from a database (MySQL) based on the identified problem. In this step, the input is information about the identified problem and the output is the solution retrieved from the database.
[1140] Step 7:
[1141] The server generates a video tutorial for solving the problem based on the solution to the identified problem. The generated solution includes a text explanation and a video link. The input is the searched solution, and the output is the video tutorial and its link.
[1142] Step 8:
[1143] The server sends the generated solution explanation and video link to the user's device. The input is the video tutorial and link, and the output is the data sent to the user's device.
[1144] Step 9:
[1145] The device displays the received solution explanation and video link to the user. Specifically, when the user clicks on the video link, the explanatory video is played. The input is the received data, and the output is the solution explanation and video link displayed to the user.
[1146] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1147] The present invention is a system that identifies problems users encounter when using services, taking into account the user's emotional state, and provides solutions. Because the user does not need to explain the problem in detail, support efficiency can be significantly improved. A specific embodiment of this system is described below.
[1148] Overall flow
[1149] 1. Start a video call
[1150] A user launches a supporting app on their device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone, generates a video call stream, and sends it to the server. The server receives the connection request and creates a new video call session.
[1151] 2. Video and audio capture
[1152] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server, which stores the received video and audio streams in a buffer and prepares them for analysis.
[1153] 3. Data Analysis
[1154] The server extracts each frame from the received video stream and uses an image analysis engine to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine and analyzes what the user is saying. The results of these analyses are combined to identify the problem the user is facing.
[1155] 4. Emotion analysis
[1156] The server uses an emotion engine that recognizes the user's emotions from the video and audio data, for example, facial expressions and tone of voice to determine whether the user is feeling impatient, angry, sad, etc.
[1157] 5. Solution Generation
[1158] The server searches for the best solution from the database based on the identified problem and the user's emotional state. If necessary, it adjusts the priority and presentation of the solutions taking the emotional state into account. It also generates an appropriate video tutorial for solving the problem and creates an explanatory video. It also generates a text description and video link for the solution.
[1159] 6. Providing solutions
[1160] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and if the user clicks on the video, the device launches a video player and plays the explanatory video.
[1161] Specific examples
[1162] Example 1: If you cannot log in
[1163] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[1164] Emotion analysis: The server determines from video and audio data that the user is feeling anxious, and determines that a solution needs to be offered quickly.
[1165] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[1166] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[1167] Example 2: When an error message appears
[1168] The user uses the camera to capture the situation where an error message is displayed and speaks. The device then sends this to the server.
[1169] Sentiment analysis: The server determines from video and audio data that the user is angry, and determines that a solution should be provided in a calm and easy-to-understand format.
[1170] The server extracts the specific error message from the video data and analyzes the user's utterances from the audio data. From these results, the server determines that the "specific error message" is the cause.
[1171] The server searches the database for a solution to this error message, generates an explanatory video showing the specific solution method, and sends the generated solution explanation and video link to the user, who then displays it on the device.
[1172] As described above, the present invention provides a system that quickly and accurately identifies the problems a user is facing, taking into account the user's emotional state, and provides appropriate solutions, thereby significantly improving user satisfaction.
[1173] The processing flow will be explained below.
[1174] Step 1:
[1175] User: Launch the support app and click the "Start Video Call" button.
[1176] Step 2:
[1177] Device: Asks the user for permission to use the camera and microphone. If permission is granted, turns on the camera and microphone.
[1178] Step 3:
[1179] Terminal: Generates the video call stream, encodes the video and audio data in real time, and sends it to the server.
[1180] Step 4:
[1181] Server: Receives connection requests from users' devices and creates new video call sessions. It buffers the received video and audio streams.
[1182] Step 5:
[1183] Server: Extracts each frame from the video stream and uses an image analysis engine to identify GUI states and error messages.
[1184] Step 6:
[1185] Server: The audio stream is converted into text using a speech analysis engine, and the content of what the user is saying is analyzed.
[1186] Step 7:
[1187] Server: Integrates the results of video and audio analysis and applies rule-based filters to identify the issues the user is facing.
[1188] Step 8:
[1189] Server: Uses an emotion engine to recognize the user's emotions from video and audio data. Emotions are analyzed from the user's facial expressions and tone of voice.
[1190] Step 9:
[1191] Server: Searches for an appropriate solution from a database based on the identified problem and the user's emotional state.
[1192] Step 10:
[1193] Server: Tailors solutions to user emotions and prioritizes them as needed. Generates problem-solving video tutorials and creates instructional videos.
[1194] Step 11:
[1195] Server: Generates a text description of the solution and a video link and sends them to the user's device.
[1196] Step 12:
[1197] On your device: Display the received text description and video link to the user.
[1198] Step 13:
[1199] User: Clicks on the video link and watches a video that explains how to solve the problem.
[1200] Step 14:
[1201] Device: Plays instructional videos and helps users solve problems by following instructions.
[1202] The above steps create a system that can quickly and accurately resolve problems users encounter while using the service. The introduction of an emotion engine makes it possible to provide optimal support based on the user's emotional state.
[1203] Example 2
[1204] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1205] In today's world, users face a wide variety of problems when using services. Support for resolving these problems typically requires users to explain the problem in detail. However, if users are unable to clearly express their problems, solving the problem increases the time and effort required. Furthermore, presenting solutions without considering the user's emotional state can detract from the user experience. To address these challenges, a system is needed that can quickly and accurately identify problems and provide optimal solutions, taking the user's emotions into account.
[1206] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1207] In this invention, the server includes means for analyzing the user's emotions from the received video data and audio data, means for generating an appropriate solution based on the identified problem and the user's emotional state, and means for transmitting an explanation of the generated solution and a video link to the user's terminal. This makes it possible to quickly and accurately identify the problem and provide an optimal solution while taking into account the user's emotions, without the user having to clearly explain the problem.
[1208] "User" means a person who uses the System to request support for the Service.
[1209] "Video call" refers to a real-time means of video and audio communication between a user and a support system.
[1210] A "terminal" is a device that a user uses to interact with the support system, such as a smartphone or computer.
[1211] A "camera" is a device for capturing images that is built into or attached to a terminal.
[1212] A "microphone" is a device for capturing sound that is built into a terminal or attached externally.
[1213] "Server" is a central processing system that analyzes data received through the video call and provides appropriate solutions to the user.
[1214] "Video data" refers to image information captured by a camera.
[1215] "Audio Data" means audio information captured by a microphone.
[1216] "Analysis" is the act of extracting and processing information based on video and audio data.
[1217] "Emotion analysis" is the process of identifying a user's emotional state based on information such as facial expressions and tone of voice.
[1218] A "solution" is a specific procedure or method provided to solve a problem a user is facing.
[1219] A "video link" is a hyperlink that indicates where to watch an explanatory video, etc.
[1220] "Display" is the act of visualizing information such as text or video links on a device screen.
[1221] A "video player" is software or an application that plays an explanatory video when you click on a video link.
[1222] MODE FOR CARRYING OUT THE INVENTION
[1223] The present invention is a system for identifying problems users encounter when using services, taking into account the user's emotional state, and providing solutions, allowing users to receive prompt and accurate support without having to explain their problems in detail.
[1224] Hardware and software used
[1225] Hardware
[1226] Device: The smartphone or computer used by the user.
[1227] Camera: A built-in or external video capture device.
[1228] Microphone: Built-in or external audio capture device.
[1229] Server: A central processing system that processes and analyzes data.
[1230] software
[1231] Supporting App: An application that allows users to initiate video calls.
[1232] Image analysis engine: For example, OpenCV is used to analyze video data.
[1233] Speech analysis engine: Converts voice data into text using, for example, the Google Speech-to-Text API.
[1234] Sentiment analysis engine: Analyze the user's emotional state using, for example, the Affectiva SDK.
[1235] System action
[1236] Starting a video call
[1237] 1. The user launches the supporting app on their device and clicks the "Start Video Call" button.
[1238] 2. The device requests permission to use the camera and microphone, and if the user grants permission, it activates the camera and microphone, generates a video call stream, and sends it to the server.
[1239] 3. The server receives the connection request and creates a new video call session.
[1240] Video and audio capture
[1241] 1. The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server.
[1242] 2. The server buffers the received video and audio streams and prepares them for analysis.
[1243] Data analysis
[1244] 1. The server extracts each frame from the received video stream and uses an image analysis engine to identify the status of the graphical user interface and error messages.
[1245] 2. At the same time, the audio stream is converted into text by a speech analysis engine and the content of what the user is saying is analyzed.
[1246] 3. Synthesize the analysis results to identify the problems users are facing.
[1247] Emotion analysis
[1248] 1. The server uses an emotion analysis engine to recognize the user's emotions from video and audio data.
[1249] 2. For example, by looking at facial expressions and tone of voice, it can identify whether the user is feeling anxious, angry, sad, or other emotions.
[1250] Solution Generation
[1251] 1. The server searches the database for the best solution based on the identified problem and the user's emotional state.
[1252] 2. Adjust the priorities and presentation of solutions to take into account emotional states.
[1253] 3. Generate appropriate video tutorials and create instructional videos to solve problems.
[1254] 4. Create a text description and video link of the generated solution.
[1255] Providing solutions
[1256] 1. The server sends the generated solution explanation and video link to the user's device.
[1257] 2. The device displays the received text description and video link to the user.
[1258] 3. When the user clicks on the video link, the device launches a video player and plays the explanatory video.
[1259] Specific examples
[1260] Example 1: If you cannot log in
[1261] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[1262] Emotion analysis: The server determines from video and audio data that the user is feeling anxious, and determines that a solution needs to be offered quickly.
[1263] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data, thereby identifying the problem of "I can't log in because the password field is blank."
[1264] The server searches the database for a solution to the correct login method, generates a message saying "Password field must be filled in" and a video tutorial showing the login procedure, and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[1265] Prompt Sentence Examples
[1266] "Users report issues via video call, you analyze them including sentiment analysis, and provide solutions with video tutorials. Users say 'I can't log in.'"
[1267] The above is an example of the present invention. This system enables the user to quickly and accurately identify problems and provide optimal solutions, taking into account emotions, without the user having to clearly explain the problem.
[1268] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1269] Specific explanation of processing steps
[1270] Step 1:
[1271] Starting a video call
[1272] Input: The user launches a supporting app and clicks the "Start Video Call" button.
[1273] How it works: A user launches a supporting app and clicks the button for a video call.
[1274] Output: The device displays a dialog to allow use of the camera and microphone.
[1275] Example of how it works:
[1276] When the user clicks the "Start Video Call" button, a dialog box will appear requesting permission to use the device's camera and microphone. If the user clicks "Allow," the camera and microphone will be activated.
[1277] Step 2:
[1278] Video and audio capture
[1279] Input: Camera and microphone activated.
[1280] How it works: The device captures video with its camera and audio with its microphone.
[1281] Output: Encoded video and audio data is generated and sent to the server.
[1282] Example of how it works:
[1283] The user's face and voice are captured in real time by a camera and microphone, encoded, and then streamed to a server.
[1284] Step 3:
[1285] Receiving and buffering data
[1286] Input: Video and audio data sent to the server.
[1287] Operation: The server stores the received video and audio data in a temporary buffer and prepares it for analysis.
[1288] Output: The data stored in the buffer is passed to the analysis engine.
[1289] Example of how it works:
[1290] When the server receives the video and audio streaming data, it is temporarily stored in a buffer, through which the data is passed to the next analysis step.
[1291] Step 4:
[1292] Data analysis
[1293] Input: Buffered video and audio data.
[1294] How it works: The server uses an image analysis engine (e.g., OpenCV) to analyze each frame of video data and recognize graphical user interfaces and error messages, while simultaneously using a speech analysis engine (e.g., Google Speech-to-Text API) to convert the audio data into text and analyze what the user is saying.
[1295] Output: Text data of the identified problem and what the user said.
[1296] Example of how it works:
[1297] A specific error message is read from the video data, and the statement "I can't log in" is converted into text from the audio data.
[1298] Step 5:
[1299] Emotion analysis
[1300] Input: Buffered video and audio data, and the identified issue.
[1301] How it works: The server uses an emotion analysis engine (e.g., Affectiva SDK) to analyze the user's emotional state, for example, by identifying the user's emotions (annoyance, anger, sadness, etc.) from facial expressions and tone of voice.
[1302] Output: Analysis data of the user's emotional state.
[1303] Example of how it works:
[1304] A video of the user's impatient face after seeing the identified error message is analyzed to identify the emotion of impatience.
[1305] Step 6:
[1306] Solution Generation
[1307] Input: Analysis data of identified problems and emotional states.
[1308] How it works: The server searches for the best solution from a database based on the identified problem and the user's emotional state, adjusts the priority and presentation of solutions as needed, and generates a video tutorial for solving the problem and creates an instructional video.
[1309] Output: A text description of the solution and a video link.
[1310] Example of how it works:
[1311] The server generates a "Password field is required" message and a video tutorial showing login steps.
[1312] Step 7:
[1313] Providing solutions
[1314] Input: A text description of the solution and a video link.
[1315] Operation: The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and when the user clicks the video link, it launches a video player and plays the explanatory video.
[1316] Output: The solution shown to the user and a video link.
[1317] Example of how it works:
[1318] The device screen displays the message "Password field requires input" along with a video link. When the user clicks the link, a video player launches and a video with the solution is played.
[1319] (Application example 2)
[1320] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1321] When users encounter security-related problems, conventional support systems often delay problem identification and solution provision, making it difficult to provide a prompt and appropriate response to the user. Furthermore, users who are in an emotional state, such as frustration, anger, or sadness, may find it difficult to understand the support content. This tends to reduce user satisfaction and prolong the problem. Therefore, there is a need for a system that not only speeds up problem identification but also provides solutions that take into account the user's emotional state.
[1322] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing received video and audio data to identify a problem faced by the user, means for analyzing the user's emotional state from the received video and audio data, and means for generating a solution based on the identified problem and the user's emotional state. This makes it possible, when a user faces a security-related problem, to quickly identify the problem and provide an effective solution that takes the user's emotional state into consideration.
[1323] "Video calling" is a function that allows users to communicate remotely in real time using audio and video.
[1324] A "terminal" is a device that a user uses to initiate a video call and capture video and audio. Examples include a smartphone or tablet.
[1325] "Server" is a central processing unit that analyzes video and audio data received from users, identifies problems, and provides solutions.
[1326] "Video data" refers to video information captured by a camera, including the user's face and the state of the screen interface.
[1327] "Audio data" is acoustic information captured by a microphone, including the user's voice and what is being said.
[1328] "Analysis" is the process by which the server extracts useful information from the data it receives and understands its meaning.
[1329] "Emotional state" refers to the emotions that a user expresses through video and audio data, such as impatience, anger, sadness, etc.
[1330] A "solution" is a method, procedure, or supporting information provided by the server to solve an identified problem.
[1331] "Description and Video Link" is a textual summary of the solution and access to a tutorial video related to that solution.
[1332] "Database" means an information management system that accumulates and stores solutions and related information in a searchable format.
[1333] "Video tutorials" are video contents that visually explain to users how to solve problems.
[1334] Overall flow
[1335] The invention includes a system for quickly and effectively responding when a user experiences a security-related problem, which takes into account the user's emotional state to identify the problem and provide a solution.
[1336] Hardware and Software
[1337] Hardware used:
[1338] Smartphone or tablet (e.g. iPhone, Android device)
[1339] Server (high-performance server for sensing and analysis, such as Amazon Web Services)
[1340] Software used:
[1341] Video calling libraries (e.g. WebRTC)
[1342] Sentiment analysis engine (e.g. Microsoft Azure Emotion API)
[1343] Speech analysis engine (e.g. Google Cloud Speech-to-Text)
[1344] A database (e.g., MySQL or PostgreSQL)
[1345] Video platform (e.g. YouTube API)
[1346] System Operation
[1347] 1. Start a video call
[1348] A user launches a specific supporting application on a smartphone or tablet. This application uses WebRTC to provide video calling functionality. When the user clicks the "Start Video Call" button, the camera and microphone are activated and video and audio capture begins.
[1349] 2. Capture of video and audio data
[1350] Video data captured by the smartphone camera and audio data captured by the microphone are encoded in real time and sent to a server, which stores the data in a buffer and prepares it for analysis.
[1351] 3. Data Analysis
[1352] The server divides the received video stream into frames and identifies the face and screen interface state. Libraries such as OpenCV are used for image analysis. Additionally, audio data is converted into text using Google Cloud Speech-to-Text, allowing the user's speech to be analyzed.
[1353] 4. Emotion analysis
[1354] The server uses the received video and audio data to determine the user's emotional state. It uses the Microsoft Azure Emotion API to analyze emotions from facial expressions and complements emotions from the user's tone of voice.
[1355] 5. Solution Generation
[1356] The server searches for an appropriate solution from a database based on the identified problem and the user's emotional state, and then generates a problem-solving explanation text and related tutorial video using the YouTube API.
[1357] 6. Providing solutions
[1358] The server sends the generated solution explanation and video link to the user's device, which displays it to the user. When the user clicks on the video link, a video player is launched and the explanatory video is played.
[1359] Specific examples
[1360] Sentiment analysis prompt example:
[1361] Identify the user's emotions from their facial expressions and tone of voice, specifically whether they are anger, sadness, confusion, or impatience.
[1362] Solution search prompt example:
[1363] Find a solution to the security issue based on the problem the user is facing and their emotional state. Describe your recommended solution and steps, and provide links to relevant video tutorials.
[1364] Thus, an embodiment of the invention provides a system that quickly and accurately solves security problems while taking into account the emotional state of the user.
[1365] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1366] Step 1:
[1367] A user launches a supporting application and clicks the "Start Video Call" button. This receives user input and activates the camera and microphone as output. This prepares the application for capturing video and audio. Specifically, the application uses the WebRTC library to request permission to use the camera and microphone.
[1368] Step 2:
[1369] The device begins capturing video data with the camera and audio data with the microphone. The captured video and audio input data is encoded and sent to the server in real time. The specific operation is a process that uses WebRTC to generate video and audio streams and send the data to the server. The output is the video and audio data sent to the server.
[1370] Step 3:
[1371] The server splits the video stream into frames and uses an image analysis engine (e.g., OpenCV) to determine the GUI state and error messages. This is the process of analyzing the captured video data (input) and extracting specific screen elements and messages (output). The specific operation is to apply image processing algorithms and identify important elements.
[1372] Step 4:
[1373] The server converts the audio stream into text using Google Cloud Speech-to-Text. The input is the transmitted audio data, and the output is the analyzed text data. The specific operation is to send the audio data to the speech analysis engine and obtain the resulting text.
[1374] Step 5:
[1375] The server integrates the analysis results of the video and audio data to identify the problem the user is facing. This is the process of combining the outputs from the image analysis engine and the audio analysis engine to identify the problem. Specifically, it centrally manages the analysis results and applies an algorithm to identify the cause of the problem. The input is the image analysis results and the audio analysis results, and the output is the identified problem.
[1376] Step 6:
[1377] The server performs emotion analysis from the video and audio data. It uses the Microsoft Azure Emotion API to identify the user's emotional state. At this stage, the input is the user's video and audio data, and the output is the analyzed emotional state. The specific operation is to send the data to the emotion analysis engine and obtain the resulting emotional information.
[1378] Step 7:
[1379] The server searches for an appropriate solution from a database based on the identified problem and the user's emotional state. The input is the analysis results and the user's emotional state, and the output is an appropriate solution. The specific operation is to quickly extract a solution using a database search algorithm.
[1380] Step 8:
[1381] The server generates an explanatory video to solve the problem. It also generates links to related tutorial videos using the YouTube API. The input is the searched solution, and the output is a solution explanation including a video link. The specific operation is the process of linking with the YouTube API and extracting the appropriate video.
[1382] Step 9:
[1383] The server sends the generated solution explanation and video link to the user's device. The input is the solution explanation and video link, and the output is data transmission to the user's device. The specific operation is to send data to the user's device via network communication.
[1384] Step 10:
[1385] The device displays the received solution explanation and video link to the user. When the user clicks on the video link, a video player on the device is launched and the explanatory video is played. The input is the solution explanation and video link sent from the server, and the output is the display and video playback on the device. The specific operation is to display the solution on the user interface and launch the video player.
[1386] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1387] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1388] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1389] [Fourth embodiment]
[1390] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1391] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1392] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1393] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1394] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1395] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1396] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1397] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1398] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1399] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1400] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1401] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1402] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1403] The present invention is a system for identifying problems that users encounter when using a service and providing solutions to those problems. This system significantly improves support efficiency because it eliminates the need for users to explain the problem in detail. A specific embodiment of this system is described below.
[1404] Overall flow
[1405] 1. Start a video call
[1406] A user launches a supporting app on their device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone, generates a video call stream, and sends it to the server. The server receives the connection request and creates a new video call session.
[1407] 2. Video and audio capture
[1408] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server, which stores the received video and audio streams in a buffer and prepares them for analysis.
[1409] 3. Data Analysis
[1410] The server extracts each frame from the received video stream and uses an image analysis engine to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine and analyzes what the user is saying. The results of these analyses are combined to identify the problem the user is facing.
[1411] 4. Solution Generation
[1412] The server searches the database for the best solution for the identified problem, generates a video tutorial for solving the problem, creates an explanatory video, and generates a text description of the solution and a video link.
[1413] 5. Providing solutions
[1414] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and if the user clicks on the video, plays the explanatory video.
[1415] Specific examples
[1416] Example 1: If you cannot log in
[1417] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[1418] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[1419] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[1420] Example 2: When an error message appears
[1421] The user uses the camera to capture the situation where an error message is displayed and speaks. The device then sends this to the server.
[1422] The server extracts the specific error message from the video data and analyzes the user's utterances from the audio data. From these results, the server determines that the "specific error message" is the cause.
[1423] The server searches the database for a solution to this error message, generates an explanatory video showing the specific solution method, and sends the generated solution explanation and video link to the user, who then displays it on the device.
[1424] As described above, the present invention provides a system for quickly and accurately identifying problems faced by users and providing appropriate solutions, thereby significantly improving user satisfaction.
[1425] The processing flow will be explained below.
[1426] Step 1:
[1427] User: Launch the support app and click the "Start Video Call" button.
[1428] Step 2:
[1429] Device: Asks the user for permission to use the camera and microphone. If permission is granted, turns on the camera and microphone and generates the video call stream.
[1430] Step 3:
[1431] Terminal: Sends the video call stream to the server. Video and audio data are encoded in real time and continuously transmitted to the server.
[1432] Step 4:
[1433] Server: Receives connection requests from users' devices and creates new video call sessions. It buffers the received video and audio streams.
[1434] Step 5:
[1435] Server: Extracts each frame from the received video stream and uses an image analysis engine to identify GUI states and error messages.
[1436] Step 6:
[1437] Server: Converts the audio stream into text using a speech analysis engine. Analyzes what the user says and extracts keywords.
[1438] Step 7:
[1439] Server: Combines the results of video and audio analysis, applies rule-based filters, and identifies the problem the user is facing.
[1440] Step 8:
[1441] Server: Searches for a suitable solution from a database based on the identified problem, and generates an instructional video demonstrating the solution if necessary.
[1442] Step 9:
[1443] Server: Generates a text description of the solution and a video link and sends them to the user's device.
[1444] Step 10:
[1445] On the device: The received text description and video link are displayed to the user. If the user clicks on the video link, the video player is launched and the explanatory video is played.
[1446] Step 11:
[1447] User: Check the solution displayed on the device and take the necessary action. If the solution is insufficient, ask for additional support via video call again.
[1448] Example 1
[1449] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1450] Conventional user support systems require users to describe their problems in detail, a process that takes time and effort. Furthermore, identifying problems and providing solutions is often done manually, which is often inefficient. These issues result in lower user satisfaction and increased support costs.
[1451] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1452] In this invention, the server includes means for analyzing the status of the graphical user interface and error messages from the received video data, means for analyzing the content of user comments from the received audio data, and means for identifying the user's problem by integrating the results of the analysis of the video data and the audio data, thereby making it possible to efficiently and quickly identify the problem and provide a solution without the user having to explain the problem in detail.
[1453] The "means for initiating a video call" refers to the operation performed by a user to initiate a video call using an application on a terminal, and the combination of software and hardware that realizes this operation.
[1454] The "means for capturing and transmitting video and audio" is a mechanism for capturing video and audio data in real time using the camera and microphone of the terminal, encoding this data, and transmitting it to the server.
[1455] The "means for analyzing the received video and audio data to identify the problem the user is facing" is a software process in which the server processes the received video and audio data using an analysis engine to identify the problem or issue the user is facing.
[1456] The "means for generating a solution" is a combination of software and hardware that allows the server to search for an appropriate solution from the database and generate a solution to provide to the user based on the searched solution.
[1457] The "means for transmitting the generated solution description and video link to the user's terminal" is a communication means for transmitting the server-generated text description of the solution and the associated video link to the user's terminal via a network.
[1458] The "means for displaying the received solution description and video link to the user" is a mechanism including a user interface for displaying the text description and video link of the solution received by the terminal from the server in a format that is easy for the user to view.
[1459] "Means for analyzing the status of the graphical user interface and error messages" refers to a software process in which the server analyzes video data to extract information such as on-screen buttons, menus, and error messages, and then uses this information to identify the user's operational status and problems.
[1460] "Means for analyzing the user's speech" refers to a software process by which the server converts the speech data into text using speech recognition technology and understands from the text what issue the user is reporting.
[1461] The "means for searching for an appropriate solution from the database" is a mechanism by which the server searches for and obtains the most appropriate solution for a specified problem from among the solutions stored in the database.
[1462] The "means for generating video tutorials" refers to video generation technology and software that allows the server to create videos that include specific steps to solve a user's problem.
[1463] The "means for transmitting to the user's terminal" refers to a communication means for transmitting the video tutorial and solution information generated by the server to the user's terminal via a network.
[1464] The "means for playing the video tutorial to the user" refers to a mechanism including a media player that allows the user to play and watch the video tutorial received by the user's terminal.
[1465] MODE FOR CARRYING OUT THE INVENTION
[1466] The present invention is a system for identifying problems users encounter when using a service and providing solutions to those problems. This system significantly improves support efficiency because it does not require users to describe the problem in detail.
[1467] Starting a video call
[1468] When a user launches a supporting app and clicks the "Start Video Call" button, the device requests permission to use the camera and microphone. If permission is granted, the device activates the camera and microphone, generates a video call stream, and sends it to the server. This stream uses technologies such as WebRTC. The server receives the connection request and creates a new video call session.
[1469] Video and audio capture
[1470] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server. The encoding is done using H.264 and AAC encoders. The server stores the received video and audio streams in a buffer and prepares them for analysis.
[1471] Data analysis
[1472] The server extracts each frame from the received video stream and uses an image analysis engine (e.g., OpenCV or TensorFlow) to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine (e.g., Google Speech-to-Text or IBM Watson Speech to Text) and analyzes what the user is saying. The results of this analysis are combined to identify the problem the user is facing.
[1473] Solution Generation
[1474] The server searches for the best solution for the identified problem from a database (e.g., MySQL or MongoDB). It then generates a video tutorial to solve the problem. Adobe After Effects and FFmpeg are used to generate the video. It then generates a text description of the solution and a video link.
[1475] Providing solutions
[1476] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user. When the user clicks the video link, the device plays the explanatory video.
[1477] Specific examples
[1478] Example 1: If you cannot log in
[1479] The user launches the support app and says, "I can't log in." The device uses its camera and microphone to send video and audio data to the server. The server identifies the blank fields on the login screen from the video data and analyzes the audio data. This allows the server to identify the problem: "I can't log in because the password field is blank." The server searches its database for a solution to the correct login method and generates a video tutorial showing the login procedure along with the message, "An entry is required in the password field." This is sent to the user's device, which then displays an explanation of the solution and a video link.
[1480] Example 2: When an error message appears
[1481] The user films the situation in which "an error message is displayed" and speaks about it. The device sends this to the server. The server extracts the specific error message from the video data and analyzes the user's statement from the audio data. As a result, the server determines that the "specific error message" is the cause. The server searches a database for a solution to this error message and generates an explanatory video showing a specific method of dealing with the problem. The generated solution explanation and video link are sent to the user, and the device displays them to the user.
[1482] Example prompts for generative AI models
[1483] Below is an example of a prompt sentence to input to the generative AI model.
[1484] 1. "A user initiates a video call. Video and audio are captured and sent to the server."
[1485] 2. "The server analyzes the video data received and checks for error messages and the status of the GUI."
[1486] 3. "Based on the identified problem, we generate optimal solutions and tutorial videos."
[1487] As described above, this system is able to quickly and accurately identify user problems and provide appropriate solutions, which is expected to improve user satisfaction and significantly improve support efficiency.
[1488] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1489] Step 1:
[1490] Starting a video call
[1491] The user launches the support app and clicks the "Start Video Call" button.
[1492] Input: User action.
[1493] The device will display a popup requesting permission to use the camera and microphone, and if the user allows it, the camera and microphone will be activated.
[1494] Input: User permission.
[1495] The device generates a video call stream and sends it to the server using WebRTC technology. The device sets the stream's protocol information and starts the stream.
[1496] Output: Video call stream.
[1497] The server receives the connection request and creates a new video call session. The server generates a session ID and sends a confirmation message to the device.
[1498] Output: A new video call session.
[1499] Step 2:
[1500] Video and audio capture
[1501] The device uses a camera to capture video data and a microphone to capture audio data.
[1502] Input: Video and audio data.
[1503] The captured video and audio data is encoded in real time using H.264 and AAC encoders.
[1504] The terminal transmits the encoded video and audio data to the server.
[1505] Output: Encoded video and audio data.
[1506] The server stores the received data in a buffer and prepares it for analysis. The server manages the buffer memory and checks the data integrity.
[1507] Output: The data stored in the buffer.
[1508] Step 3:
[1509] Data analysis
[1510] The server extracts each frame from the received video stream using a video analysis engine (e.g., OpenCV).
[1511] Input: Buffered video data.
[1512] The server uses an image analysis engine to identify GUI states (e.g., buttons, input fields, error messages).
[1513] Input: Video data for each frame.
[1514] Output: Data of the identified GUI element.
[1515] The server converts the audio stream into text using a speech analysis engine (e.g., Google Speech-to-Text), which is then analyzed by a natural language processing engine.
[1516] Input: Buffered audio data.
[1517] Output: The audio data converted to text.
[1518] The server combines the results of the video and audio analysis to identify the problem the user is facing.
[1519] Input: Parsed GUI data and text data.
[1520] Output: Identified issues.
[1521] Step 4:
[1522] Solution Generation
[1523] The server searches for a solution to the identified problem from a database that stores information about existing solutions.
[1524] Input: Identified problem.
[1525] Output: The solution found.
[1526] The server uses a video generation engine (e.g. Adobe After Effects) to generate a video tutorial for solving the problem. The video is automatically generated based on the steps you specify.
[1527] Input: The solution found.
[1528] Output: The generated video tutorial.
[1529] The server generates a message containing a text description of the solution and a video link.
[1530] Input: Retrieved solutions and generated video tutorials.
[1531] Output: Solution description and video link.
[1532] Step 5:
[1533] Providing solutions
[1534] The server generates a solution explanation and sends a video link to the user's device.
[1535] Input: Solution description and video link.
[1536] Output: Data sent to the user's terminal.
[1537] The device displays the received solution explanation and video link to the user.
[1538] Input: Data received from the server.
[1539] Output: The solution description and video link that is displayed to the user.
[1540] When the user clicks on the video link, the device will play the instructional video.
[1541] Input: Click action on the video link.
[1542] Output: The instruction video that will be played.
[1543] (Application example 1)
[1544] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1545] When a machine or robot operating in a factory begins to operate abnormally, it is necessary to quickly and accurately identify the problem and provide an appropriate solution. However, currently, it takes a lot of time for operators to identify the problem and find a solution. In addition, finding an appropriate solution requires specialized knowledge, which leads to a decrease in production efficiency. Therefore, there is a need to introduce a system that can detect abnormalities in real time, efficiently identify the problem, and provide a solution.
[1546] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1547] In this invention, the server includes a means for detecting abnormal operation of the machine, a means for transmitting data captured by the machine to the server, and a means for identifying an optimal solution for the machine and instructing the machine on how to deal with the problem. This makes it possible to quickly detect abnormal operation, efficiently identify the problem, and provide an appropriate solution.
[1548] "Means for a user to initiate a video call" refers to software or hardware functionality that allows a user to initiate a video call session to receive remote support.
[1549] "Means for the terminal to capture and transmit the user's video and audio" refers to a function that enables the terminal to use a camera and microphone to capture the user's video and audio and transmit them to a server.
[1550] "Means for analyzing the video and audio data received by the server to identify the problem the user is facing" is a function for analyzing the video and audio data received by the server and identifying the problem the user is facing.
[1551] The "means by which the server generates a solution to the identified problem" is a function by which the server generates an appropriate solution to the problem after identifying the problem.
[1552] "Means for transmitting the explanation of the solution generated by the server and a video link to the user's terminal" is a function for transmitting the explanation of the solution generated by the server and a related video link to the user's terminal.
[1553] The "means for displaying to the user the explanation of the solution and the video link received by the terminal" is a function for displaying to the user the explanation of the solution and the video link received by the user's terminal.
[1554] "Means for machines to detect abnormal operation" refers to software or hardware functions that allow machines operating in a factory to monitor their operating status and detect any abnormalities that may occur.
[1555] "Means for transmitting data captured by the machine to the server" is a function for transmitting video and audio data captured by the machine to the server when an abnormality is detected.
[1556] "Means for identifying the optimal solution for the machine and instructing the machine on how to respond" is a function that enables the server to analyze a problem, identify an appropriate solution, and instruct the machine on how to respond.
[1557] MODE FOR CARRYING OUT THE INVENTION
[1558] A system for implementing this invention includes various means for quickly and accurately identifying a user's problem and providing an appropriate solution. Specifically, a user initiates a video call, and the video and audio data captured by the terminal is sent to a server, which analyzes the data and identifies the problem. The server then generates a solution and provides it to the user. In a factory, if a machine detects an abnormal operation, it notifies the server, which then analyzes the situation and provides an optimal solution.
[1559] Hardware and Software
[1560] Factory robots: equipped with cameras, microphones, and communication modules, including sensors to detect abnormal behavior.
[1561] Server: A video analysis engine (e.g., OpenCV), an audio analysis engine (e.g., Google Speech-to-Text API), a database (e.g., MySQL), and a solution generation module (Python script).
[1562] Device: Includes the device used by the user (smartphone, tablet, PC, etc.).
[1563] Data processing and calculation
[1564] The server first receives video and audio data sent from terminals and factory robots. The received video data is analyzed using an image analysis engine (OpenCV) to identify GUI status and error messages. Meanwhile, the audio data is converted into text using a speech analysis engine (Google Speech-to-Text API), and the user's speech is analyzed.
[1565] The analyzed data is integrated and the server uses this information to identify the problems the user is facing. For the identified problems, the server searches the database (MySQL) for the best solution and generates a text explanation and a video tutorial as the solution.
[1566] The generated solution is sent from the server to the user's device, where it is displayed to the user, along with notifications and explanatory video links sent from the device.
[1567] Specific examples
[1568] 1. Resolving login issues:
[1569] The user says, "I can't log in."
[1570] The device uses a camera and microphone to capture video and audio and transmits them to a server.
[1571] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[1572] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays the solution explanation and video link to the user.
[1573] Example prompt sentence:
[1574] Identify issues from video call data and provide solutions.
[1575] Audio data: "I can't log in"
[1576] Video data: Screenshot of the login screen
[1577] 2. Machine Anomaly Detection and Resolution:
[1578] A factory machine detects abnormal operation.
[1579] The machine uses a camera and microphone to capture video and audio and transmits it to a server.
[1580] The server identifies specific error messages and abnormal behavior patterns from the video data and also analyzes information from the audio data. Based on the results, the server identifies the cause of the abnormality and searches a database for an appropriate solution.
[1581] The server instructs the machine on the generated solutions and workarounds, allowing it to make the necessary corrections.
[1582] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1583] Step 1:
[1584] The user launches the supporting app on the device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone. The input obtained here is the user's video and audio, which are generated as a video call stream and sent to the server.
[1585] Step 2:
[1586] The server buffers the video call stream received from the device. The received video data is divided into frames, and the audio data is encoded. The input to this analysis preparation stage is the video and audio streams, and the output is data ready for analysis.
[1587] Step 3:
[1588] The server uses an image analysis engine (OpenCV) to analyze each frame of the received video stream. Specifically, it performs processing to identify the GUI status and error messages. The input is the received video data, and the output is information about the GUI status and error messages.
[1589] Step 4:
[1590] The server converts the received audio stream into text using a speech analysis engine (Google Speech-to-Text API). This converts the user's input voice data into text data and analyzes the user's speech. The input is voice data, and the output is text data.
[1591] Step 5:
[1592] The server integrates the results of the video data analysis obtained in step 3 and the results of the audio data analysis obtained in step 4. This identifies the problem the user is facing. The inputs are the video analysis results and the audio analysis results, and the output is information about the identified problem.
[1593] Step 6:
[1594] The server searches for a suitable solution from a database (MySQL) based on the identified problem. In this step, the input is information about the identified problem and the output is the solution retrieved from the database.
[1595] Step 7:
[1596] The server generates a video tutorial for solving the problem based on the solution to the identified problem. The generated solution includes a text explanation and a video link. The input is the searched solution, and the output is the video tutorial and its link.
[1597] Step 8:
[1598] The server sends the generated solution explanation and video link to the user's device. The input is the video tutorial and link, and the output is the data sent to the user's device.
[1599] Step 9:
[1600] The device displays the received solution explanation and video link to the user. Specifically, when the user clicks on the video link, the explanatory video is played. The input is the received data, and the output is the solution explanation and video link displayed to the user.
[1601] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1602] The present invention is a system that identifies problems users encounter when using services, taking into account the user's emotional state, and provides solutions. Because the user does not need to explain the problem in detail, support efficiency can be significantly improved. A specific embodiment of this system is described below.
[1603] Overall flow
[1604] 1. Start a video call
[1605] A user launches a supporting app on their device and clicks the "Start Video Call" button. The device requests permission to use the camera and microphone, and if permission is granted, activates the camera and microphone, generates a video call stream, and sends it to the server. The server receives the connection request and creates a new video call session.
[1606] 2. Video and audio capture
[1607] The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server, which stores the received video and audio streams in a buffer and prepares them for analysis.
[1608] 3. Data Analysis
[1609] The server extracts each frame from the received video stream and uses an image analysis engine to identify the GUI status and error messages. At the same time, it converts the audio stream into text using a speech analysis engine and analyzes what the user is saying. The results of these analyses are combined to identify the problem the user is facing.
[1610] 4. Emotion analysis
[1611] The server uses an emotion engine that recognizes the user's emotions from the video and audio data, for example, facial expressions and tone of voice to determine whether the user is feeling impatient, angry, sad, etc.
[1612] 5. Solution Generation
[1613] The server searches for the best solution from the database based on the identified problem and the user's emotional state. If necessary, it adjusts the priority and presentation of the solutions taking the emotional state into account. It also generates an appropriate video tutorial for solving the problem and creates an explanatory video. It also generates a text description and video link for the solution.
[1614] 6. Providing solutions
[1615] The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and if the user clicks on the video, the device launches a video player and plays the explanatory video.
[1616] Specific examples
[1617] Example 1: If you cannot log in
[1618] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[1619] Emotion analysis: The server determines from video and audio data that the user is feeling anxious, and determines that a solution needs to be offered quickly.
[1620] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data. From these results, the server identifies the problem as "I can't log in because the password field is blank."
[1621] The server searches the database for a solution on the correct login method, generates a video tutorial showing the login procedure along with the message "Password field is required", and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[1622] Example 2: When an error message appears
[1623] The user uses the camera to capture the situation where an error message is displayed and speaks. The device then sends this to the server.
[1624] Sentiment analysis: The server determines from video and audio data that the user is angry, and determines that a solution should be provided in a calm and easy-to-understand format.
[1625] The server extracts the specific error message from the video data and analyzes the user's utterances from the audio data. From these results, the server determines that the "specific error message" is the cause.
[1626] The server searches the database for a solution to this error message, generates an explanatory video showing the specific solution method, and sends the generated solution explanation and video link to the user, who then displays it on the device.
[1627] As described above, the present invention provides a system that quickly and accurately identifies the problems a user is facing, taking into account the user's emotional state, and provides appropriate solutions, thereby significantly improving user satisfaction.
[1628] The processing flow will be explained below.
[1629] Step 1:
[1630] User: Launch the support app and click the "Start Video Call" button.
[1631] Step 2:
[1632] Device: Asks the user for permission to use the camera and microphone. If permission is granted, turns on the camera and microphone.
[1633] Step 3:
[1634] Terminal: Generates the video call stream, encodes the video and audio data in real time, and sends it to the server.
[1635] Step 4:
[1636] Server: Receives connection requests from users' devices and creates new video call sessions. It buffers the received video and audio streams.
[1637] Step 5:
[1638] Server: Extracts each frame from the video stream and uses an image analysis engine to identify GUI states and error messages.
[1639] Step 6:
[1640] Server: The audio stream is converted into text using a speech analysis engine, and the content of what the user is saying is analyzed.
[1641] Step 7:
[1642] Server: Integrates the results of video and audio analysis and applies rule-based filters to identify the issues the user is facing.
[1643] Step 8:
[1644] Server: Uses an emotion engine to recognize the user's emotions from video and audio data. Emotions are analyzed from the user's facial expressions and tone of voice.
[1645] Step 9:
[1646] Server: Searches for an appropriate solution from a database based on the identified problem and the user's emotional state.
[1647] Step 10:
[1648] Server: Tailors solutions to user emotions and prioritizes them as needed. Generates problem-solving video tutorials and creates instructional videos.
[1649] Step 11:
[1650] Server: Generates a text description of the solution and a video link and sends them to the user's device.
[1651] Step 12:
[1652] On your device: Display the received text description and video link to the user.
[1653] Step 13:
[1654] User: Clicks on the video link and watches a video that explains how to solve the problem.
[1655] Step 14:
[1656] Device: Plays instructional videos and helps users solve problems by following instructions.
[1657] The above steps create a system that can quickly and accurately resolve problems users encounter while using the service. The introduction of an emotion engine makes it possible to provide optimal support based on the user's emotional state.
[1658] Example 2
[1659] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1660] In today's world, users face a wide variety of problems when using services. Support for resolving these problems typically requires users to explain the problem in detail. However, if users are unable to clearly express their problems, solving the problem increases the time and effort required. Furthermore, presenting solutions without considering the user's emotional state can detract from the user experience. To address these challenges, a system is needed that can quickly and accurately identify problems and provide optimal solutions, taking the user's emotions into account.
[1661] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1662] In this invention, the server includes means for analyzing the user's emotions from the received video data and audio data, means for generating an appropriate solution based on the identified problem and the user's emotional state, and means for transmitting an explanation of the generated solution and a video link to the user's terminal. This makes it possible to quickly and accurately identify the problem and provide an optimal solution while taking into account the user's emotions, without the user having to clearly explain the problem.
[1663] "User" means a person who uses the System to request support for the Service.
[1664] "Video call" refers to a real-time means of video and audio communication between a user and a support system.
[1665] A "terminal" is a device that a user uses to interact with the support system, such as a smartphone or computer.
[1666] A "camera" is a device for capturing images that is built into or attached to a terminal.
[1667] A "microphone" is a device for capturing sound that is built into a terminal or attached externally.
[1668] "Server" is a central processing system that analyzes data received through the video call and provides appropriate solutions to the user.
[1669] "Video data" refers to image information captured by a camera.
[1670] "Audio Data" means audio information captured by a microphone.
[1671] "Analysis" is the act of extracting and processing information based on video and audio data.
[1672] "Emotion analysis" is the process of identifying a user's emotional state based on information such as facial expressions and tone of voice.
[1673] A "solution" is a specific procedure or method provided to solve a problem a user is facing.
[1674] A "video link" is a hyperlink that indicates where to watch an explanatory video, etc.
[1675] "Display" is the act of visualizing information such as text or video links on a device screen.
[1676] A "video player" is software or an application that plays an explanatory video when you click on a video link.
[1677] MODE FOR CARRYING OUT THE INVENTION
[1678] The present invention is a system for identifying problems users encounter when using services, taking into account the user's emotional state, and providing solutions, allowing users to receive prompt and accurate support without having to explain their problems in detail.
[1679] Hardware and software used
[1680] Hardware
[1681] Device: The smartphone or computer used by the user.
[1682] Camera: A built-in or external video capture device.
[1683] Microphone: Built-in or external audio capture device.
[1684] Server: A central processing system that processes and analyzes data.
[1685] software
[1686] Supporting App: An application that allows users to initiate video calls.
[1687] Image analysis engine: For example, OpenCV is used to analyze video data.
[1688] Speech analysis engine: Converts voice data into text using, for example, the Google Speech-to-Text API.
[1689] Sentiment analysis engine: Analyze the user's emotional state using, for example, the Affectiva SDK.
[1690] System action
[1691] Starting a video call
[1692] 1. The user launches the supporting app on their device and clicks the "Start Video Call" button.
[1693] 2. The device requests permission to use the camera and microphone, and if the user grants permission, it activates the camera and microphone, generates a video call stream, and sends it to the server.
[1694] 3. The server receives the connection request and creates a new video call session.
[1695] Video and audio capture
[1696] 1. The device encodes the video data captured by the camera and the audio data captured by the microphone in real time and sends them to the server.
[1697] 2. The server buffers the received video and audio streams and prepares them for analysis.
[1698] Data analysis
[1699] 1. The server extracts each frame from the received video stream and uses an image analysis engine to identify the status of the graphical user interface and error messages.
[1700] 2. At the same time, the audio stream is converted into text by a speech analysis engine and the content of what the user is saying is analyzed.
[1701] 3. Synthesize the analysis results to identify the problems users are facing.
[1702] Emotion analysis
[1703] 1. The server uses an emotion analysis engine to recognize the user's emotions from video and audio data.
[1704] 2. For example, by looking at facial expressions and tone of voice, it can identify whether the user is feeling anxious, angry, sad, or other emotions.
[1705] Solution Generation
[1706] 1. The server searches the database for the best solution based on the identified problem and the user's emotional state.
[1707] 2. Adjust the priorities and presentation of solutions to take into account emotional states.
[1708] 3. Generate appropriate video tutorials and create instructional videos to solve problems.
[1709] 4. Create a text description and video link of the generated solution.
[1710] Providing solutions
[1711] 1. The server sends the generated solution explanation and video link to the user's device.
[1712] 2. The device displays the received text description and video link to the user.
[1713] 3. When the user clicks on the video link, the device launches a video player and plays the explanatory video.
[1714] Specific examples
[1715] Example 1: If you cannot log in
[1716] The user launches the support app and says, "I can't log in." The device uses the camera and microphone to send video and audio data to the server.
[1717] Emotion analysis: The server determines from video and audio data that the user is feeling anxious, and determines that a solution needs to be offered quickly.
[1718] The server identifies the blank fields on the login screen from the video data and analyzes the utterance "I can't log in" from the audio data, thereby identifying the problem of "I can't log in because the password field is blank."
[1719] The server searches the database for a solution to the correct login method, generates a message saying "Password field must be filled in" and a video tutorial showing the login procedure, and sends it to the user's device, which then displays an explanation of the solution and a video link to the user.
[1720] Prompt Sentence Examples
[1721] "Users report issues via video call, you analyze them including sentiment analysis, and provide solutions with video tutorials. Users say 'I can't log in.'"
[1722] The above is an example of the present invention. This system enables the user to quickly and accurately identify problems and provide optimal solutions, taking into account emotions, without the user having to clearly explain the problem.
[1723] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1724] Specific explanation of processing steps
[1725] Step 1:
[1726] Starting a video call
[1727] Input: The user launches a supporting app and clicks the "Start Video Call" button.
[1728] How it works: A user launches a supporting app and clicks the button for a video call.
[1729] Output: The device displays a dialog to allow use of the camera and microphone.
[1730] Example of how it works:
[1731] When the user clicks the "Start Video Call" button, a dialog box will appear requesting permission to use the device's camera and microphone. If the user clicks "Allow," the camera and microphone will be activated.
[1732] Step 2:
[1733] Video and audio capture
[1734] Input: Camera and microphone activated.
[1735] How it works: The device captures video with its camera and audio with its microphone.
[1736] Output: Encoded video and audio data is generated and sent to the server.
[1737] Example of how it works:
[1738] The user's face and voice are captured in real time by a camera and microphone, encoded, and then streamed to a server.
[1739] Step 3:
[1740] Receiving and buffering data
[1741] Input: Video and audio data sent to the server.
[1742] Operation: The server stores the received video and audio data in a temporary buffer and prepares it for analysis.
[1743] Output: The data stored in the buffer is passed to the analysis engine.
[1744] Example of how it works:
[1745] When the server receives the video and audio streaming data, it is temporarily stored in a buffer, through which the data is passed to the next analysis step.
[1746] Step 4:
[1747] Data analysis
[1748] Input: Buffered video and audio data.
[1749] How it works: The server uses an image analysis engine (e.g., OpenCV) to analyze each frame of video data and recognize graphical user interfaces and error messages, while simultaneously using a speech analysis engine (e.g., Google Speech-to-Text API) to convert the audio data into text and analyze what the user is saying.
[1750] Output: Text data of the identified problem and what the user said.
[1751] Example of how it works:
[1752] A specific error message is read from the video data, and the statement "I can't log in" is converted into text from the audio data.
[1753] Step 5:
[1754] Emotion analysis
[1755] Input: Buffered video and audio data, and the identified issue.
[1756] How it works: The server uses an emotion analysis engine (e.g., Affectiva SDK) to analyze the user's emotional state, for example, by identifying the user's emotions (annoyance, anger, sadness, etc.) from facial expressions and tone of voice.
[1757] Output: Analysis data of the user's emotional state.
[1758] Example of how it works:
[1759] A video of the user's impatient face after seeing the identified error message is analyzed to identify the emotion of impatience.
[1760] Step 6:
[1761] Solution Generation
[1762] Input: Analysis data of identified problems and emotional states.
[1763] How it works: The server searches for the best solution from a database based on the identified problem and the user's emotional state, adjusts the priority and presentation of solutions as needed, and generates a video tutorial for solving the problem and creates an instructional video.
[1764] Output: A text description of the solution and a video link.
[1765] Example of how it works:
[1766] The server generates a "Password field is required" message and a video tutorial showing login steps.
[1767] Step 7:
[1768] Providing solutions
[1769] Input: A text description of the solution and a video link.
[1770] Operation: The server sends the generated solution explanation and video link to the user's device. The device displays the received text explanation and video link to the user, and when the user clicks the video link, it launches a video player and plays the explanatory video.
[1771] Output: The solution shown to the user and a video link.
[1772] Example of how it works:
[1773] The device screen displays the message "Password field requires input" along with a video link. When the user clicks the link, a video player launches and a video with the solution is played.
[1774] (Application example 2)
[1775] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1776] When users encounter security-related problems, conventional support systems often delay problem identification and solution provision, making it difficult to provide a prompt and appropriate response to the user. Furthermore, users who are in an emotional state, such as frustration, anger, or sadness, may find it difficult to understand the support content. This tends to reduce user satisfaction and prolong the problem. Therefore, there is a need for a system that not only speeds up problem identification but also provides solutions that take into account the user's emotional state.
[1777] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing received video and audio data to identify a problem faced by the user, means for analyzing the user's emotional state from the received video and audio data, and means for generating a solution based on the identified problem and the user's emotional state. This makes it possible, when a user faces a security-related problem, to quickly identify the problem and provide an effective solution that takes the user's emotional state into consideration.
[1778] "Video calling" is a function that allows users to communicate remotely in real time using audio and video.
[1779] A "terminal" is a device that a user uses to initiate a video call and capture video and audio. Examples include a smartphone or tablet.
[1780] "Server" is a central processing unit that analyzes video and audio data received from users, identifies problems, and provides solutions.
[1781] "Video data" refers to video information captured by a camera, including the user's face and the state of the screen interface.
[1782] "Audio data" is acoustic information captured by a microphone, including the user's voice and what is being said.
[1783] "Analysis" is the process by which the server extracts useful information from the data it receives and understands its meaning.
[1784] "Emotional state" refers to the emotions that a user expresses through video and audio data, such as impatience, anger, sadness, etc.
[1785] A "solution" is a method, procedure, or supporting information provided by the server to solve an identified problem.
[1786] "Description and Video Link" is a textual summary of the solution and access to a tutorial video related to that solution.
[1787] "Database" means an information management system that accumulates and stores solutions and related information in a searchable format.
[1788] "Video tutorials" are video contents that visually explain to users how to solve problems.
[1789] Overall flow
[1790] The invention includes a system for quickly and effectively responding when a user experiences a security-related problem, which takes into account the user's emotional state to identify the problem and provide a solution.
[1791] Hardware and Software
[1792] Hardware used:
[1793] Smartphone or tablet (e.g. iPhone, Android device)
[1794] Server (high-performance server for sensing and analysis, such as Amazon Web Services)
[1795] Software used:
[1796] Video calling libraries (e.g. WebRTC)
[1797] Sentiment analysis engine (e.g. Microsoft Azure Emotion API)
[1798] Speech analysis engine (e.g. Google Cloud Speech-to-Text)
[1799] A database (e.g., MySQL or PostgreSQL)
[1800] Video platform (e.g. YouTube API)
[1801] System Operation
[1802] 1. Start a video call
[1803] A user launches a specific supporting application on a smartphone or tablet. This application uses WebRTC to provide video calling functionality. When the user clicks the "Start Video Call" button, the camera and microphone are activated and video and audio capture begins.
[1804] 2. Capture of video and audio data
[1805] Video data captured by the smartphone camera and audio data captured by the microphone are encoded in real time and sent to a server, which stores the data in a buffer and prepares it for analysis.
[1806] 3. Data Analysis
[1807] The server divides the received video stream into frames and identifies the face and screen interface state. Libraries such as OpenCV are used for image analysis. Additionally, audio data is converted into text using Google Cloud Speech-to-Text, allowing the user's speech to be analyzed.
[1808] 4. Emotion analysis
[1809] The server uses the received video and audio data to determine the user's emotional state. It uses the Microsoft Azure Emotion API to analyze emotions from facial expressions and complements emotions from the user's tone of voice.
[1810] 5. Solution Generation
[1811] The server searches for an appropriate solution from a database based on the identified problem and the user's emotional state, and then generates a problem-solving explanation text and related tutorial video using the YouTube API.
[1812] 6. Providing solutions
[1813] The server sends the generated solution explanation and video link to the user's device, which displays it to the user. When the user clicks on the video link, a video player is launched and the explanatory video is played.
[1814] Specific examples
[1815] Sentiment analysis prompt example:
[1816] Identify the user's emotions from their facial expressions and tone of voice, specifically whether they are anger, sadness, confusion, or impatience.
[1817] Solution search prompt example:
[1818] Find a solution to the security issue based on the problem the user is facing and their emotional state. Describe your recommended solution and steps, and provide links to relevant video tutorials.
[1819] Thus, an embodiment of the invention provides a system that quickly and accurately solves security problems while taking into account the emotional state of the user.
[1820] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1821] Step 1:
[1822] A user launches a supporting application and clicks the "Start Video Call" button. This receives user input and activates the camera and microphone as output. This prepares the application for capturing video and audio. Specifically, the application uses the WebRTC library to request permission to use the camera and microphone.
[1823] Step 2:
[1824] The device begins capturing video data with the camera and audio data with the microphone. The captured video and audio input data is encoded and sent to the server in real time. The specific operation is a process that uses WebRTC to generate video and audio streams and send the data to the server. The output is the video and audio data sent to the server.
[1825] Step 3:
[1826] The server splits the video stream into frames and uses an image analysis engine (e.g., OpenCV) to determine the GUI state and error messages. This is the process of analyzing the captured video data (input) and extracting specific screen elements and messages (output). The specific operation is to apply image processing algorithms and identify important elements.
[1827] Step 4:
[1828] The server converts the audio stream into text using Google Cloud Speech-to-Text. The input is the transmitted audio data, and the output is the analyzed text data. The specific operation is to send the audio data to the speech analysis engine and obtain the resulting text.
[1829] Step 5:
[1830] The server integrates the analysis results of the video and audio data to identify the problem the user is facing. This is the process of combining the outputs from the image analysis engine and the audio analysis engine to identify the problem. Specifically, it centrally manages the analysis results and applies an algorithm to identify the cause of the problem. The input is the image analysis results and the audio analysis results, and the output is the identified problem.
[1831] Step 6:
[1832] The server performs emotion analysis from the video and audio data. It uses the Microsoft Azure Emotion API to identify the user's emotional state. At this stage, the input is the user's video and audio data, and the output is the analyzed emotional state. The specific operation is to send the data to the emotion analysis engine and obtain the resulting emotional information.
[1833] Step 7:
[1834] The server searches for an appropriate solution from a database based on the identified problem and the user's emotional state. The input is the analysis results and the user's emotional state, and the output is an appropriate solution. The specific operation is to quickly extract a solution using a database search algorithm.
[1835] Step 8:
[1836] The server generates an explanatory video to solve the problem. It also generates links to related tutorial videos using the YouTube API. The input is the searched solution, and the output is a solution explanation including a video link. The specific operation is the process of linking with the YouTube API and extracting the appropriate video.
[1837] Step 9:
[1838] The server sends the generated solution explanation and video link to the user's device. The input is the solution explanation and video link, and the output is data transmission to the user's device. The specific operation is to send data to the user's device via network communication.
[1839] Step 10:
[1840] The device displays the received solution explanation and video link to the user. When the user clicks on the video link, a video player on the device is launched and the explanatory video is played. The input is the solution explanation and video link sent from the server, and the output is the display and video playback on the device. The specific operation is to display the solution on the user interface and launch the video player.
[1841] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1842] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1843] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1844] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1845] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1846] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1847] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1848] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1849] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1850] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1851] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1852] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1853] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1854] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1855] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1856] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1857] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1858] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1859] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1860] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1861] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1862] The following is further disclosed regarding the above embodiment.
[1863] (Claim 1)
[1864] means for a user to initiate a video call;
[1865] means for the terminal to capture and transmit video and audio of the user;
[1866] means for analyzing the video and audio data received by the server to identify a problem the user is experiencing;
[1867] means for the server to generate a solution to the identified problem;
[1868] A means for the server to send the generated solution explanation and video link to the user's terminal;
[1869] a means for displaying to the user the solution explanation and video link received by the terminal;
[1870] A system including:
[1871] (Claim 2)
[1872] A means to analyze the GUI status and error messages from the video data received by the server,
[1873] A means for analyzing the content of a user's speech from the voice data received by the server;
[1874] a means for the server to integrate the analysis results of the video data and the audio data to identify a problem of the user;
[1875] 10. The system of claim 1, comprising:
[1876] (Claim 3)
[1877] a means for the server to search a database for an appropriate solution based on the identified problem;
[1878] a means for the server to generate a video tutorial for solving the problem;
[1879] A means for the server to transmit the generated video tutorial to a user's terminal;
[1880] means for playing the received video tutorial to the user;
[1881] 10. The system of claim 1, comprising:
[1882] "Example 1"
[1883] (Claim 1)
[1884] means for a user to initiate a video call;
[1885] means for the terminal to capture and transmit video and audio of the user;
[1886] means for analyzing the video and audio data received by the server to identify a problem the user is experiencing;
[1887] means for the server to generate a solution to the identified problem;
[1888] A means for the server to send the generated solution explanation and video link to the user's terminal;
[1889] a means for displaying to the user the solution explanation and video link received by the terminal;
[1890] A system including:
[1891] (Claim 2)
[1892] A means for analyzing the status of a graphical user interface and error messages from the video data received by the server;
[1893] A means for analyzing the content of a user's speech from the voice data received by the server;
[1894] a means for the server to integrate the analysis results of the video data and the audio data to identify a problem of the user;
[1895] 10. The system of claim 1, comprising:
[1896] (Claim 3)
[1897] a means for the server to search a database for an appropriate solution based on the identified problem;
[1898] a means for the server to generate a video tutorial for solving the problem;
[1899] A means for the server to transmit the generated video tutorial to a user's terminal;
[1900] means for playing the received video tutorial to the user;
[1901] 10. The system of claim 1, comprising:
[1902] "Application Example 1"
[1903] (Claim 1)
[1904] means for a user to initiate a video call;
[1905] means for the terminal to capture and transmit video and audio of the user;
[1906] means for analyzing the video and audio data received by the server to identify a problem the user is experiencing;
[1907] means for the server to generate a solution to the identified problem;
[1908] A means for the server to send the generated solution explanation and video link to the user's terminal;
[1909] a means for displaying to the user the solution explanation and video link received by the terminal;
[1910] means for the machine to detect abnormal operation;
[1911] A system including:
[1912] (Claim 2)
[1913] A means to analyze the GUI status and error messages from the video data received by the server,
[1914] A means for analyzing the content of a user's speech from the voice data received by the server;
[1915] a means for the server to integrate the analysis results of the video data and the audio data to identify a problem of the user;
[1916] means for transmitting the machine-captured data to a server;
[1917] 10. The system of claim 1, comprising:
[1918] (Claim 3)
[1919] a means for the server to search a database for an appropriate solution based on the identified problem;
[1920] a means for the server to generate a video tutorial for solving the problem;
[1921] A means for the server to transmit the generated video tutorial to a user's terminal;
[1922] means for playing the received video tutorial to the user;
[1923] A means of identifying the best solution for the machine and instructing the machine on how to respond;
[1924] 10. The system of claim 1, comprising:
[1925] "Example 2: Combining Emotion Engines"
[1926] (Claim 1)
[1927] means for a user to initiate a video call;
[1928] means for the terminal to capture and transmit video and audio of the user;
[1929] means for analyzing the video and audio data received by the server to identify a problem the user is experiencing;
[1930] A server analyzes user emotions from video data and audio data;
[1931] means for the server to generate an appropriate solution based on the identified problem and the user's emotional state;
[1932] A means for the server to send the generated solution explanation and video link to the user's terminal;
[1933] a means for displaying to the user the solution explanation and video link received by the terminal;
[1934] A device detects when a user clicks on a video link, and then starts a video player to play the explanatory video;
[1935] A system including:
[1936] (Claim 2)
[1937] A means for analyzing the status of a graphical user interface and error messages from the video data received by the server;
[1938] A means for analyzing the content of a user's speech from the voice data received by the server;
[1939] a means for the server to integrate the analysis results of the video data and the audio data to identify a problem of the user;
[1940] A server analyzes the emotional state of a user from the video data and audio data;
[1941] 10. The system of claim 1, comprising:
[1942] (Claim 3)
[1943] means for the server to search a database for an appropriate solution based on the identified problem and the user's emotional state;
[1944] a means for the server to generate a video tutorial for solving the problem;
[1945] A means for the server to transmit the generated video tutorial to a user's terminal;
[1946] means for playing the received video tutorial to the user;
[1947] 10. The system of claim 1, comprising:
[1948] "Application example 2 when combining emotion engines"
[1949] (Claim 1)
[1950] means for a user to initiate a video call;
[1951] means for the terminal to capture and transmit video and audio of the user;
[1952] means for analyzing the video and audio data received by the server to identify a problem the user is experiencing;
[1953] means for analyzing the emotional state of the user from the video and audio data received by the server;
[1954] means for the server to generate a solution based on the identified problem and the user's emotional state;
[1955] A means for the server to send the generated solution explanation and video link to the user's terminal;
[1956] a means for displaying to the user the solution explanation and video link received by the terminal;
[1957] A system including:
[1958] (Claim 2)
[1959] A means for analyzing the screen interface status and error messages from the video data received by the server;
[1960] A means for analyzing the content of a user's speech from the voice data received by the server;
[1961] a means for the server to integrate the analysis results of the video data and the audio data to identify a problem of the user;
[1962] a means of adjusting solution priorities based on emotional state;
[1963] 10. The system of claim 1, comprising:
[1964] (Claim 3)
[1965] a means for the server to search a database for an appropriate solution based on the identified problem;
[1966] a means for the server to generate a video tutorial for solving the problem;
[1967] A means for the server to transmit the generated video tutorial to a user's terminal;
[1968] means for playing the received video tutorial to the user;
[1969] A means for generating an explanatory video according to the content of the error message;
[1970] 10. The system of claim 1, comprising: [Explanation of symbols]
[1971] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for a user to initiate a video call; means for the terminal to capture and transmit video and audio of the user; means for analyzing the video and audio data received by the server to identify a problem the user is experiencing; means for the server to generate a solution to the identified problem; A means for the server to send the generated solution explanation and video link to the user's terminal; a means for displaying to the user the solution explanation and video link received by the terminal; A system including:
2. A means to analyze the GUI status and error messages from the video data received by the server, A means for analyzing the content of a user's speech from the voice data received by the server; a means for the server to integrate the analysis results of the video data and the audio data to identify a problem of the user; The system of claim 1 , comprising:
3. a means for the server to search a database for an appropriate solution based on the identified problem; a means for the server to generate a video tutorial for solving the problem; A means for the server to transmit the generated video tutorial to a user's terminal; means for playing the received video tutorial to the user; The system of claim 1 , comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A