system
The system addresses the inefficiencies in self-training by recognizing voice input, converting it to text, and providing relevant information via voice, thereby improving the efficiency and reducing excessive advice in practice fields.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-02
- Publication Date
- 2026-04-14
AI Technical Summary
Existing systems fail to quickly and accurately provide specific information to users during self-training, particularly in practice fields, leading to inefficiencies and excessive advice, which hampers the learning process.
A system that recognizes voice input, converts it to text, analyzes the intent, sends a request to a server for relevant information, and notifies the user via voice, minimizing excessive advice and improving self-learning efficiency.
Enables users to quickly and accurately obtain necessary information, enhancing the efficiency of independent practice and reducing the impact of excessive guidance.
Smart Images

Figure 2026064837000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] When a user conducts self-training, there is a problem that it is difficult to quickly and accurately obtain information about specific technologies or methods. In particular, when it is desired to check a specific part of a designated YouTube (registered trademark) video, it is necessary to repeatedly review the video, which is time-consuming and laborious. Also, it is not uncommon to be affected by excessive advice (teaching devil) from others in a practice field. As a result, the efficiency of individual practice decreases, and the effect of self-study is limited.
Means for Solving the Problems
[0005] This invention provides a system that recognizes voice input from a user and provides appropriate video information. Specifically, the system includes means for recognizing voice input from a user and converting it into text data, means for analyzing the text data to determine the user's intent, means for sending a request to a server to obtain relevant information based on the user's intent, and means for notifying the user of the information received from the server by voice. This invention enables users to quickly and accurately obtain information about specific technologies and methods, improving the efficiency of their practice. Furthermore, it allows them to proceed with self-learning without receiving excessive advice at the practice field.
[0006] "Voice input" refers to instructions or questions in voice format that a user makes through a voice input device such as a microphone.
[0007] "Speech recognition" is a technology that analyzes voice input as digital data and converts it into corresponding text data.
[0008] "Text data" refers to character data that corresponds to voice input, converted using speech recognition technology.
[0009] "User intent" refers to the content or information that the user is trying to convey through voice input.
[0010] "Analysis" is the process of analyzing text data to identify the user's intent.
[0011] A "request" is a communication message that asks a server to retrieve information.
[0012] A "server" is a computer system that receives information via a network and returns the processing results.
[0013] "Video information" refers to data such as the video title, URL, and related sections obtained from video sharing platforms.
[0014] "Notification" refers to the act of informing users of processing results and information, and here it mainly refers to being performed in voice form.
[0015] "Video sharing platform" refers to a service on the Internet where users can share and view videos.
[0016] "Instructor nag" is a common name referring to a person who gives excessive advice or guidance to others in a practice field or the like.
[0017] "Self-learning" is a process of independently learning and practicing without relying on external instructors or advisors.
Brief Explanation of Drawings
[0018] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0019] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0022] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0026] [First Embodiment]
[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0039] This invention is a system for users to quickly and accurately acquire information about specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice (overly helpful advice) at practice facilities.
[0040] User voice input
[0041] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[0042] Speech-to-text conversion and analysis
[0043] The device uses speech recognition technology to convert recorded audio into text. This converted text data is used to analyze the user's intent. Natural language processing technology is used for the analysis to identify which part of which video the user wants to know about.
[0044] Request to the server
[0045] The device sends an appropriate request to the server based on the analyzed text data. This request includes information about the video or content specified by the user. For example, a request is sent to the server to re-identify a specific YouTube video and specify the "final part" of it.
[0046] Server Processing
[0047] Based on the received request, the server retrieves relevant video information from the video sharing platform. Using existing platforms such as the YouTube API, it obtains the video's title, URL, and the content of the specified portion. The server also analyzes the video's subtitles and script to identify the information the user is seeking.
[0048] Information response
[0049] The server sends the retrieved information back to the terminal. This information includes the video title, URL, and relevant content. For example, it might provide the URL of a video titled "Basic Swing Practice" and the information that "the last part explained the key points of the follow-through."
[0050] Notification to the user
[0051] The terminal notifies the user of information received from the server via voice. This information is conveyed to the user using speech synthesis technology. As a result, the user can instantly obtain information about specific techniques and methods, and easily check the instructions needed during practice. Furthermore, it enables the user to practice independently without receiving excessive advice at the practice field.
[0052] Specific example
[0053] For example, if a user says, "Show me the last part of the swing practice video I watched yesterday by a YouTuber," the device transcribes the voice command into text and sends a request to the server based on the analysis. The server retrieves information from the YouTube video, identifies the relevant section, and sends it back to the device. The device then relays this information to the user via voice, allowing the user to improve the efficiency of their practice.
[0054] The system of this invention allows users to improve the quality of their independent practice and quickly obtain necessary information. This improves practice efficiency and makes it possible to avoid excessive intervention from others.
[0055] The following describes the processing flow.
[0056] Step 1:
[0057] The user gives voice commands using an earphone microphone. For example, they input specific instructions in voice format, such as, "Show me the last part of the swing practice video I watched yesterday again."
[0058] Step 2:
[0059] The device receives and records voice input. This recorded voice data is used for subsequent processing.
[0060] Step 3:
[0061] The device uses speech recognition technology to convert recorded audio data into text data. This is done using Google's (registered trademark) speech recognition API, among others.
[0062] Step 4:
[0063] The device analyzes what the user wants based on the converted text data. Using natural language processing technology, it extracts information to identify a specific video and its corresponding section.
[0064] Step 5:
[0065] The device sends a request to the server based on the analysis results. The request includes the URL of the video specified by the user and information about a specific part of it.
[0066] Step 6:
[0067] The server receives the request from the terminal and uses the video sharing platform's API to retrieve the relevant video information. Specifically, it retrieves the video's title, URL, and relevant content.
[0068] Step 7:
[0069] The server further analyzes the acquired video information and uses subtitles and scripts to identify what the user is looking for. For example, it might extract specific information such as "the follow-through points explained at the end."
[0070] Step 8:
[0071] The server sends the analyzed information back to the terminal. The information sent back includes the video title, URL, and a detailed description of the relevant section.
[0072] Step 9:
[0073] The terminal constructs a message to notify the user based on information received from the server. This message is then conveyed to the user in voice format using speech synthesis technology.
[0074] Step 10:
[0075] The user receives an audio notification from their device and obtains information about a specific part of the video they were looking for. Based on this information, the user can effectively proceed with their self-practice.
[0076] (Example 1)
[0077] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0078] In conventional self-practice methods, it was difficult for users to quickly and accurately obtain information on specific techniques and methods, and there was a lack of ways to avoid excessive advice, especially at the practice field. As a result, practice efficiency decreased, and it often took a long time for users to obtain the information they needed. The present invention aims to solve these problems and provide a system that allows users to improve the quality of their self-practice.
[0079] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0080] In this invention, the server includes means for recognizing voice input from a user and converting it into text data, means for analyzing the text data to determine the user's intent, means for sending a request to the server to obtain relevant information based on the user's intent, and means for notifying the user of the information received from the server by voice. This allows the user to quickly and accurately obtain the necessary information, improve the efficiency of self-practice, and minimize the impact of excessive advice.
[0081] "Voice input" refers to voice data containing instructions or information spoken by the user.
[0082] "Text data" refers to digital information obtained by converting voice input into a string format.
[0083] "Speech recognition technology" is a technology that analyzes speech information and converts its content into text data.
[0084] "Natural language processing technology" refers to techniques that analyze text data to understand the user's intentions and emotions.
[0085] A "request" is a request sent to a server to obtain specific information.
[0086] A "server" is a computer system that receives requests via a network and provides information.
[0087] A "video sharing platform" is a web service that allows users to upload, share, and watch video content.
[0088] "Speech synthesis technology" is a technology that converts text data into speech.
[0089] "Analysis results" refer to the user's intentions and necessary information obtained through the analysis of text data.
[0090] "Video information" refers to digital information about specific video content obtained from video sharing platforms.
[0091] Modes for carrying out the invention
[0092] This invention is a system for users to quickly and accurately acquire information on specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice at the practice field. Specific embodiments are described below.
[0093] User voice input
[0094] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[0095] Speech-to-text conversion
[0096] The device converts recorded audio into text using speech recognition technologies such as the Google Cloud Speech-to-Text API. This converted text data is then used to analyze the user's intent.
[0097] Text analysis
[0098] The device analyzes the converted text data using natural language processing technologies such as the Google Cloud Natural Language API. This identifies which part of which video the user wants to know about.
[0099] Sending a request to the server
[0100] The device sends an appropriate request to the server based on the analyzed text data. This request includes information about the video or content specified by the user. For example, it might include a request to specify a particular part from a particular video sharing platform.
[0101] Server Processing
[0102] Based on the received request, the server retrieves relevant video information using the YouTube API and other tools. It obtains the video title, URL, and the content of the specified section, and analyzes the subtitles and script related to that video to identify the information the user is seeking.
[0103] Information response
[0104] The server sends the retrieved information back to the terminal. This information includes the video title, URL, and relevant content. For example, it might provide the URL of a video titled "Basic Swing Practice" and the information that "the last part explained the key points of the follow-through."
[0105] Notification to the user
[0106] The device uses speech synthesis technology, such as the Google Cloud Text-to-Speech API, to notify the user of information received from the server via voice. This allows the user to instantly obtain information about specific techniques and methods, and easily check the instructions needed during practice. It also enables the user to practice independently without receiving excessive advice at the practice field.
[0107] Specific example
[0108] For example, if a user requests to see the last part of the swing practice video from the YouTuber they watched yesterday, the following process will occur:
[0109] 1. The user gives voice commands using an earphone microphone.
[0110] 2. The device records voice commands and converts them into text using speech recognition technology.
[0111] 3. The terminal parses the text and sends a request to the server.
[0112] 4. The server uses the YouTube API to retrieve the relevant video information and analyzes the content of a specific section.
[0113] 5. The server sends the retrieved information (e.g., video URL, relevant content) back to the device.
[0114] 6. The device uses speech synthesis technology to notify the user of information.
[0115] Example of a prompt
[0116] "Show me the last part of the swing practice video I saw from the YouTuber yesterday again."
[0117] This invention allows users to quickly obtain information useful for independent practice and to continue practicing efficiently.
[0118] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0119] Step 1: Recording voice input
[0120] The user transmits voice input to the device via an earphone microphone. The input voice is recorded on the device and saved as digital audio data. A specific example of this action would be the user saying, "Show me the last part of the swing practice video I watched yesterday again."
[0121] Step 2: Text conversion using speech recognition
[0122] The device sends the recorded audio data to the Google Cloud Speech-to-Text API, which converts the audio into text data. The input is the recorded audio data, and the output is the corresponding string (e.g., "Show me the last part of the swing practice video I watched yesterday again"). Specifically, this process involves analyzing the audio waveform and generating the corresponding string.
[0123] Step 3: Analyzing text data
[0124] The device sends the acquired text data to the Google Cloud Natural Language API for analysis of the user's intent. The input is the converted text data, and the output is the analysis result, including the user's intent and request. Specifically, the NLP engine analyzes the context, keywords, and intent, and recognizes it as a request regarding the final part of a particular video.
[0125] Step 4: Generate and send the request
[0126] The terminal generates and sends a request to the server to retrieve relevant information based on the analysis results. The input is the analysis results, including the user's intent, and the output is the request data sent to the server. In a specific example, a request is generated that includes "a special video URL and section for playing the final part of the swing practice."
[0127] Step 5: Collecting video information
[0128] The server processes the received request and uses the YouTube API to collect specific video information. The input is the request data sent from the device, and the output is information including the title, URL, and content of a specific part of the relevant video. Specifically, the server accesses the YouTube API to retrieve the metadata of the video and the script or subtitles of the target section.
[0129] Step 6: Return the information
[0130] The server sends the collected video information back to the device. The input is video information obtained from the YouTube API, and the output is the information sent to the device. In a specific example, data is returned that includes the URL of a video titled "Basic Swing Practice" and information about "Follow-through points explained in the last part."
[0131] Step 7: Notify the user
[0132] The device uses text-to-speech technology (Google Cloud Text-to-Speech API) to convert information received from the server into speech and notifies the user. The input is video information sent back from the server, and the output is audio data played for the user. Specifically, the device tells the user, "In the last part of the swing practice video you watched yesterday, the key points of the follow-through were explained."
[0133] Through the steps described above, the system of the present invention can quickly and accurately provide relevant video information based on the user's voice instructions. This processing flow allows the user to effectively practice independently while avoiding excessive advice.
[0134] (Application Example 1)
[0135] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0136] The operation and maintenance procedures for machinery in factories are specialized, and employees are required to acquire this information quickly and accurately. However, finding and learning the necessary information on the job is time-consuming and laborious, and can sometimes require excessive instruction. Therefore, there is a need for systems that support efficient work and self-learning among employees.
[0137] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0138] In this invention, the server includes means for recognizing voice input and converting it into text data, means for analyzing the text data to determine the user's intent, means for sending a request to an information processing device to acquire relevant information based on the user's intent, means for notifying the user of the information received from the information processing device by voice, and means for acquiring information regarding how to operate and maintain machinery in factory work. This enables employees to efficiently acquire necessary information and engage in self-learning.
[0139] "Voice input" refers to a method in which users provide information such as commands or questions using their voice via a headset microphone or similar device.
[0140] "Text data" refers to data obtained by converting voice input into characters, and is a string of characters in a format that can be analyzed by a computer.
[0141] "Analysis" is the process of analyzing text data to understand the user's intentions and requirements.
[0142] "User intent" refers to what the user is seeking through voice input, including the content of their requests and the information they desire.
[0143] A "request" is data sent to a server to retrieve relevant information based on the user's intent.
[0144] An "information processing device" refers to a system that manages and processes various types of information and data via a network.
[0145] "Notifying" refers to the process of conveying information obtained from a server or information processing device back to the user.
[0146] "Factory work" refers to all tasks performed in workplaces such as manufacturing, using machinery and equipment.
[0147] "Operating instructions" refer to the procedures and methods for making a machine or device work correctly.
[0148] "Maintenance procedures" refer to the procedures for maintenance, inspection, and repair performed to maintain the condition of machinery and equipment.
[0149] A "video sharing platform" refers to a web service that allows users to upload, play, and share videos over the internet.
[0150] This invention's system utilizes speech recognition and natural language processing technologies to provide information on how to operate and maintain machinery in a factory. Users can input voice information using an earphone microphone and quickly and accurately obtain relevant information based on that input.
[0151] System Configuration
[0152] The system consists of the following main components:
[0153] 1. Voice Input Device: The user gives voice commands using an earphone microphone. For example, they might input a voice command such as, "Show me the last part of the maintenance video I watched before again."
[0154] 2. Speech Recognition System: This system uses speech recognition technology (such as the Google Speech Recognition API) to convert recorded audio into text data. This system operates on the device.
[0155] 3. Contextual Analysis System: Natural language processing techniques (e.g., nltk or spaCy) are used to analyze text data and determine the user's intent. At this stage, relevant information is identified based on the analysis results.
[0156] 4. Server: Receives requests based on user intent and retrieves relevant video information from video sharing platforms (such as the YouTube API). It also generates a response containing the necessary information and sends it back to the device.
[0157] 5. Notification System: The terminal notifies the user of information received from the server via voice using speech synthesis technology (e.g., Pyttsx3).
[0158] Hardware and software
[0159] The following hardware and software are used to implement this system:
[0160] Hardware: Headset with microphone, smartphone or tablet, server.
[0161] Software: Google Speech Recognition API (speech recognition), nltk or spaCy (natural language processing), YouTube API (video information retrieval), Pyttsx3 (speech synthesis).
[0162] Data processing and data calculation
[0163] 1. Speech-to-text conversion: Convert user voice input into text data using the Google Speech Recognition API.
[0164] 2. Text Analysis: The generated text data is analyzed using natural language processing libraries such as nltk and spaCy to identify the user's intent.
[0165] 3. Information Request: Based on the user's request, a request for relevant video information is made to the YouTube API, and data such as the video title and URL is retrieved.
[0166] 4. Information Notification: Use Pyttsx3 to generate the acquired information in audio format and notify the user.
[0167] Specific example
[0168] If a user gives the voice command, "Show me the last part of the robot maintenance video I saw before again," the system will perform the following actions:
[0169] 1. Convert voice input to text.
[0170] 2. Analyze the converted text data to identify what information the user is looking for.
[0171] 3. Use the YouTube API to retrieve relevant video information.
[0172] 4. The acquired information (video title, URL, specified portion) is generated using speech synthesis technology and notified to the user.
[0173] Example of a prompt
[0174] User: "Show me the last part of the robot maintenance video I saw before again."
[0175] This allows users to efficiently obtain necessary information and streamline factory operations. The system reduces excessive on-site instruction and supports employees' self-directed learning.
[0176] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0177] Step 1:
[0178] The user gives voice commands using an earphone microphone. For example, they might input a voice command such as, "Show me the last part of the robot maintenance video I watched before again." The input is the user's voice data.
[0179] Step 2:
[0180] The device uses speech recognition technology (Google Speech Recognition API) to convert the user's voice data into text data. In this process, voice data is provided as input, and the converted text data is output.
[0181] Step 3:
[0182] The device uses natural language processing technology (nltk or spaCy) to analyze the generated text data and identify the user's intent. In this step, text data is the input, and the analysis outputs the user's intent (e.g., a specific part of a video to play).
[0183] Step 4:
[0184] The terminal sends a request to the server to retrieve relevant information based on the user's intent. In this step, data including the user's intent is input, and the request data to the server is output.
[0185] Step 5:
[0186] Based on the received request, the server retrieves relevant video information from the video sharing platform (YouTube API). In this step, the request data is used as input, and the video title, URL, and specific information are output.
[0187] Step 6:
[0188] The server sends the acquired information back to the terminal. In this step, video information is used as input and is sent to the terminal.
[0189] Step 7:
[0190] The terminal uses speech synthesis technology (Pyttsx3) to notify the user of the information received from the server via voice. In this step, the received video information is the input, and the output is an audio notification.
[0191] Step 8:
[0192] The user receives an audio notification from their device and reconfirms information on how to operate or maintain the relevant machine. In this step, the audio notification is provided as input, and an action is output that plays the information the user is requesting.
[0193] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0194] This invention is a system for users to quickly and accurately acquire information about specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice (overly helpful advice) at practice facilities. Furthermore, this invention can further enhance the user experience by incorporating an emotion engine that recognizes the user's emotions.
[0195] User voice input
[0196] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[0197] Speech-to-text conversion and emotion recognition
[0198] The device uses speech recognition technology to convert recorded speech into text. This utilizes a speech recognition API. The converted text data is used to analyze the user's intent. Simultaneously, an emotion engine is used to analyze the user's emotions from the speech. The emotion engine analyzes the tone and characteristics of the speech to determine the user's emotional state.
[0199] User intent analysis and request generation
[0200] The device analyzes the user's intent based on text data and sentiment data. Combining natural language processing and sentiment analysis technologies, it extracts information to identify a specific video and its associated sections. Based on the analysis results, it sends an appropriate request to the server. This request includes information about the video and content specified by the user.
[0201] Server Processing
[0202] Based on the request received from the terminal, the server uses the video sharing platform's API to retrieve relevant video information. Specifically, it obtains the video's title, URL, and relevant content. Furthermore, the server analyzes the subtitles and script related to the video to identify the information the user is seeking.
[0203] Information response and notification
[0204] The server sends the analyzed information back to the device. This information includes the video title, URL, and a detailed description of the relevant parts. Based on the information received from the server, the device constructs a message to notify the user. Using speech synthesis technology, this message is delivered to the user in voice format. The emotion engine adjusts the content and tone of the notification according to the user's emotional state. For example, if the user is anxious, the information is delivered in a calm tone; conversely, if the user is relaxed, the information is delivered in a more casual tone.
[0205] Specific example
[0206] For example, if a user says, "Show me the last part of the YouTuber's swing practice video I watched yesterday," the device transcribes the voice command into text and then uses an emotion engine to analyze the user's emotional state. Based on the analyzed text data and emotion data, the device identifies the user's intent and sends a request to the server. The server retrieves information from the YouTube video, identifies the relevant section, and sends it back to the device. The device then uses this information, along with the emotion engine's analysis results, to send a voice notification. For example, if the user is anxious, the notification might say, "Please listen calmly. The last part of the video about the basics of swing practice explained the key points of the follow-through. The URL is http: / / example.com."
[0207] The system of this invention allows users to improve the quality of their independent practice and quickly obtain necessary information. This improves practice efficiency and allows them to avoid excessive intervention from others. Furthermore, by considering the user's emotional state, it can provide a more human-like interaction.
[0208] The following describes the processing flow.
[0209] Step 1:
[0210] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again."
[0211] Step 2:
[0212] The device records voice input. This recorded voice data is used for speech recognition processing.
[0213] Step 3:
[0214] The device uses speech recognition technology to convert recorded audio data into text data. It utilizes a speech recognition API to accurately transcribe the user's speech into text.
[0215] Step 4:
[0216] The device analyzes the converted text data to identify the user's intent. Using natural language processing technology, it determines the video or specific part of the video the user is looking for.
[0217] Step 5:
[0218] The device simultaneously uses an emotion engine to analyze the user's emotional state from the voice data. It analyzes the tone of voice and speaking style to determine whether the user is excited, calm, or otherwise emotional.
[0219] Step 6:
[0220] The device sends an appropriate request to the server based on the analyzed text and sentiment data. This request includes information about the video identified by the user and the content of relevant parts of it.
[0221] Step 7:
[0222] Based on the request received from the terminal, the server uses the video sharing platform's API to retrieve the relevant video information. Specifically, it retrieves the video's title, URL, and relevant content.
[0223] Step 8:
[0224] The server further analyzes the acquired video information to identify detailed explanations of relevant parts. For example, it extracts the content explained in the final part of the video from the script or subtitles.
[0225] Step 9:
[0226] The server returns the analyzed information to the terminal. This information includes the video title, URL, and a detailed description of the relevant section.
[0227] Step 10:
[0228] The terminal constructs a message to notify the user based on information received from the server. The message is created using speech synthesis technology and adjusted according to the user's emotional state, which is analyzed by the emotion engine.
[0229] Step 11:
[0230] The device provides voice notifications and conveys acquired information to the user. For example, it might notify the user of the content and URL explained at the end of a video titled "Basic Swing Practice" in a calm or casual tone, depending on the user's emotional state.
[0231] Step 12:
[0232] Users receive audio notifications from their devices and obtain information about specific parts of the video they need. This allows users to effectively practice on their own.
[0233] (Example 2)
[0234] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0235] Conventional self-training and learning support systems have struggled to provide appropriate information quickly and accurately based on user voice input. Furthermore, they fail to consider the user's emotional state, resulting in a poor user experience. To solve these problems, the present invention aims to provide a system that efficiently transcribes user voice instructions into text, performs emotion analysis, and provides appropriate video information.
[0236] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0237] In this invention, the server includes means for recognizing voice input from a user and converting it into text data; means for analyzing the user's emotions from the text data and voice data; means for determining the user's intentions by analyzing the text data and emotion data; means for sending a request to the server to obtain relevant information based on the user's intentions; means for analyzing information obtained from the server through a video sharing platform and identifying the parts requested by the user; means for notifying the user of the information received from the server by voice; and means for adjusting the content and tone of the notification according to the user's emotional state. This enables the user to quickly and accurately obtain the necessary information and receive feedback that is appropriate to their emotions.
[0238] A "user" is an individual or organization that is using the system to obtain information.
[0239] "Voice input" refers to voice data used by a user to give instructions to a system via an earphone microphone or other voice input device.
[0240] "Text data" refers to character information converted from voice input using speech recognition technology.
[0241] "Emotional data" refers to information obtained by analyzing the user's emotional state from voice data.
[0242] "Speech recognition technology" is a technology that converts voice input into text data, and typically uses APIs or specialized software.
[0243] "Emotional analysis technology" is a technology that analyzes a user's emotional state from voice data, analyzing the tone of voice and characteristics of their speaking style.
[0244] "Natural language processing technology" is a technology that analyzes text data to understand the user's intent.
[0245] A "request" is a request for information to be retrieved, sent to a server based on the user's intent.
[0246] A "server" is a computer system that processes information based on requests and provides it to users.
[0247] A "video sharing platform" is a website or service that hosts videos that are accessible to users.
[0248] "Subtitles" are information used to display the content of a video in text format.
[0249] A "script" is text data that records the content and dialogue of a video.
[0250] "Speech synthesis technology" is a technology that converts text data into a speech format for playback.
[0251] "Notification content" refers to the specific information conveyed to the user based on information obtained from the server.
[0252] "Tone" refers to the emotion or attitude expressed in the audio of a notification.
[0253] "User intent" refers to the purpose or desired information that the user is trying to convey to the system through voice commands.
[0254] This invention is a system designed to enable users to quickly and accurately acquire information about specific techniques and methods when practicing independently. It improves the efficiency of independent learning by recognizing user voice input and providing appropriate video information. Furthermore, it can enhance the user experience by recognizing user emotions. This system includes three main elements: the user, the terminal, and the server.
[0255] Hardware and software to be used
[0256] 1. Earphone microphone (hardware): Used by the user to give voice commands.
[0257] 2. Device (hardware and software): A device for recording voice input, converting it to text, performing sentiment analysis, and intent analysis. The Google Cloud Speech-to-Text API is used for text data analysis, and the Azure Cognitive Services Emotion API is used for sentiment analysis.
[0258] 3. Server (hardware and software): Based on user instructions, it retrieves video information, analyzes that information, and identifies the necessary parts. The YouTube Data API is used for this purpose.
[0259] System operation
[0260] The user gives voice commands using an earphone microphone. For example, they might give specific instructions such as, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[0261] The device converts recorded audio into text data using the Google Cloud Speech-to-Text API. Simultaneously, it analyzes the user's emotional state from the audio data using the Azure Cognitive Services Emotion API. This text data and emotional data are then used for further analysis.
[0262] The device analyzes the user's intent using OpenAI's GPT-4 model based on text and sentiment data. Based on the analysis results, the specific video and its corresponding section that the user requests are identified. A request containing this information is then sent to the server.
[0263] The server uses the YouTube Data API to search for the specified video information and retrieves the video's title, URL, content, etc. The server then analyzes the video's subtitles and script to identify the parts requested by the user.
[0264] The server sends the analyzed information back to the device. The device uses the Google Cloud Text-to-Speech API to generate a voice message based on the information received from the server. At the same time, it adjusts the tone of the notification based on the sentiment analysis results. For example, if the user is anxious, the information will be delivered in a calm tone.
[0265] The device plays the generated voice message to the user. This allows the user to quickly and accurately obtain the necessary information.
[0266] Specific example
[0267] For example, if a user instructs, "Show me the last part of the swing practice video I watched yesterday again," the device records this audio and transcribes it using the Google Cloud Speech-to-Text API. Next, it performs sentiment analysis using the Azure Cognitive Services Emotion API. The analyzed text and sentiment data are further analyzed using the OpenAI GPT-4 model to identify a specific video segment. Based on this information, a request is sent to the server, and the relevant video information is retrieved using the YouTube Data API. The retrieved information is sent back to the device, and an audio message is generated using the Google Cloud Text-to-Speech API. The device then plays this message to the user.
[0268] For example, the user might prompt with the following message: "Show me the last part of the swing practice video I watched yesterday by a YouTuber." Based on this, the system operates and provides the user with the necessary video information.
[0269] This allows users to improve the quality of their self-practice and obtain necessary information quickly and accurately. Furthermore, feedback that takes into account the user's emotional state can provide a more human-like interaction.
[0270] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0271] Step 1:
[0272] The user gives voice commands using an earphone microphone. These voice commands are specific, such as, "Show me the last part of the swing practice video I watched yesterday again." The device records this voice and prepares for the next processing step.
[0273] Input: User's voice instructions
[0274] Output: Recorded audio data
[0275] Step 2:
[0276] The terminal converts the recorded audio data into text data using the Google Cloud Speech-to-Text API. The text data is the extraction of the content of the voice instruction as character information. Also, the terminal analyzes the user's emotional state from the audio data using the Emotion API of Azure Cognitive Services.
[0277]
[0278]
[0279] <Input: Recorded audio data
[0280] <Output: Text data and emotion data
[0281]
[0282]
[0283] <Step 3:
[0284] <The terminal analyzes the user's intention using the GPT-4 model of OpenAI based on the obtained text data and emotion data. Specifically, it identifies the specific part of the video that the user requests.
[0285] <Input: Text data and emotion data
[0286] <Output: Analyzed intention data (information on specific parts of the video)
[0287] <Step 4:<The terminal sends a request to the server to obtain appropriate information based on the analyzed intention data. The request to the server includes information about the video and content specified by the user. The request includes information about the title and parts of a specific video.
[0285] <Input: Intention data
[0286] <Output: Request to be sent to the server
[0287] Step 5:
[0288] Based on the request received from the terminal, the server uses the API of a video sharing platform (e.g., YouTube) to search for and retrieve the relevant video information. Specifically, it obtains the video title, URL, and relevant content, and further analyzes subtitles and scripts to identify the requested portion.
[0289] Input: Request from terminal
[0290] Output: Acquired video information (title, URL, subtitles, content of specific section)
[0291] Step 6:
[0292] The server sends the acquired video information back to the device. Based on the received information, the device generates an audio notification message using the Google Cloud Text-to-Speech API. This message includes information about the specific part of the video requested by the user. It also adjusts the tone of the notification based on the user's sentiment data.
[0293] Input: Acquired video information and emotion data
[0294] Output: Generated voice notification message
[0295] Step 7:
[0296] The device plays a generated voice notification message to the user. For example, if the user is anxious, it will notify the user in a calm tone, saying, "Please listen calmly. The last part of the video on the basics of swing practice explained the key points of the follow-through. The URL is http: / / example.com." Through voice notifications, the user can quickly and accurately obtain the information they need.
[0297] Input: Voice notification message
[0298] Output: Voice notification to the user
[0299] This allows users to improve the quality of their self-practice and quickly and accurately obtain the information they need. Furthermore, by considering the user's emotional state, a better user experience can be provided.
[0300] (Application Example 2)
[0301] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0302] While conventional learning support systems could retrieve relevant information based on user voice commands, they struggled to provide information appropriately according to the user's emotional state. This resulted in insufficient improvement in user learning efficiency and experience. Furthermore, particularly in content distribution services, there was a lack of means to quickly retrieve the learning content learners desired and to provide appropriate interactions that responded to the user's emotions.
[0303] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0304] In this invention, the server includes means for sending a request to the server to obtain relevant information based on the user's intent; means for notifying the user of the information received from the server by voice; means for determining the user's emotional state using emotion analysis technology; means for adjusting the content and tone of the notification based on the information received from the server and the user's emotional state; and means for learners to obtain learning content through voice input in a content distribution service. This enables the provision of information appropriately according to the user's emotional state, improving learning efficiency and enabling more human-like interaction.
[0305] "Voice input" refers to voice data emitted by a user using a voice collection device such as a microphone.
[0306] "Text data" refers to data obtained by converting voice input into character information using technologies such as voice recognition technology.
[0307] "Request" refers to the content of a request sent to a server to obtain relevant information based on a user's intention or requirement.
[0308] "Server" refers to a computer or computer system that has the role of storing information, performing processing, and providing information to clients via a network.
[0309] "Voice recognition technology" is a technology for analyzing voice input collected by a microphone or the like and converting it into text data.
[0310] "Video sharing platform" refers to an online service through which users can upload, share, and view video content via the Internet.
[0311] "Sentiment analysis technology" is a technology for analyzing data such as voice and text to estimate a user's emotional state.
[0312] "Notification" refers to an act in which a system transmits information or a message to a user.
[0313] "Tone" refers to the characteristics of the pitch or speaking style of a voice in a message such as a notification.
[0314] "Content delivery service" refers to a service that distributes digital content such as videos and music to users via the Internet.
[0315] "Learning content" refers to educational videos, texts, and other multimedia-formatted materials used by users for learning.
[0316] The system for implementing this invention uses the following hardware and software.
[0317] Hardware and software to be used
[0318] Microphone (sound acquisition device):
[0319] It is used to collect voice input from users.
[0320] Device (smartphone, etc.):
[0321] This device performs speech recognition and emotion analysis technologies, processes data, and notifies the user of the results.
[0322] Speech recognition API:
[0323] To convert voice input into text data, you can use, for example, Google's speech recognition API.
[0324] Generative AI models:
[0325] To perform emotion analysis, for example, we can use the pipeline from the "transformers" library.
[0326] server:
[0327] The system uses the API of a video sharing platform (e.g., the YouTube Data API) to retrieve relevant video information and sends that information back to the device.
[0328] Speech synthesis technology:
[0329] This is used to notify users of information from the server as voice messages.
[0330] Explanation of the process
[0331] 1. Collection and recognition of voice input:
[0332] The user provides voice input via a microphone. For example, they might give a command like, "Show me the last part of the swing practice video I watched yesterday again."
[0333] The device collects this audio and uses speech recognition technology to convert it into text data.
[0334] 2. Emotional analysis and intention judgment:
[0335] The device uses a generative AI model to analyze the user's emotional state from text data. In this process, it uses a generative AI model (for example, the "transformers" library's pipeline) to identify emotions.
[0336] The device uses speech recognition technology to analyze text data and determine the user's intent. For example, it can understand that the user is requesting to play a specific video.
[0337] 3. Sending requests and retrieving information:
[0338] The device sends a request to the server to retrieve relevant information based on the user's intent. This request includes the search query for the required video.
[0339] The server uses an API to retrieve relevant video information (title, URL, etc.) from the video sharing platform.
[0340] 4. Information response and notification:
[0341] The server sends the acquired video information back to the terminal.
[0342] The device adjusts the content and tone of notifications based on the information it receives and the user's emotional state.
[0343] The device uses speech synthesis technology to generate an appropriate voice message and notify the user.
[0344] Specific example
[0345] For example, if a user instructs, "Show me the part of the math lesson from yesterday where we saw the proof of Fermat's Little Theorem again":
[0346] 1. User voice collection and recognition:
[0347] The user's voice commands are collected by the device and converted into text data using speech recognition technology, such as "Show me again the part of the math lesson yesterday that showed the proof of Fermat's Little Theorem."
[0348] 2. Emotional analysis and intention judgment:
[0349] Text data is analyzed using a generation AI model to identify the user's emotional state.
[0350] The device determines that the user is requesting the playback of a specific math video.
[0351] 3. Sending requests and retrieving information:
[0352] Based on user instructions, a request is sent to the server, and relevant mathematical video information is retrieved from the video sharing platform.
[0353] 4. Information response and notification:
[0354] Based on information received from the server, the device adjusts the tone of notifications according to the user's emotional state.
[0355] For example, an audio message is generated and sent to the user saying, "Here is the video proof of Fermat's Little Theorem that you requested. The URL is http: / / example.com."
[0356] Example of a prompt
[0357] Perform sentiment analysis on the AI model using the following prompt:
[0358] "Please analyze the emotion in the following sentence: 'Show me again the part of the math lesson yesterday that showed the proof of Fermat's Little Theorem.'"
[0359] This embodiment allows users to quickly obtain the necessary learning materials and receive appropriate support tailored to their emotional state.
[0360] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0361] Step 1:
[0362] Collection of user voice input
[0363] The user uses a microphone to provide voice input. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again."
[0364] Input: User's voice data
[0365] Output: Audio data is collected.
[0366] Step 2:
[0367] Text conversion of audio data
[0368] The device uses a speech recognition API to convert audio data into text data. Specifically, it uses speech recognition technology (for example, Google's speech recognition API) to convert speech into text information.
[0369] Input: Collected audio data
[0370] Output: Text data (Example: "Show me the last part of the swing practice video I watched yesterday again.")
[0371] Step 3:
[0372] Emotion analysis
[0373] The device uses a generative AI model to analyze the user's emotional state from text data. It uses emotion analysis techniques (e.g., the transformers library's pipeline) to analyze the tone and characteristics of speech to determine the emotional state.
[0374] Input: Text data
[0375] Output: Emotion analysis results (e.g., "anxiety")
[0376] Step 4:
[0377] Intent analysis and request generation
[0378] The device analyzes the user's intent based on text data and sentiment analysis results. Specifically, it uses natural language processing technology to identify what the user is looking for.
[0379] Based on the analysis results, the terminal sends a request to the server to retrieve relevant information.
[0380] Input: Text data, sentiment analysis results
[0381] Output: Request (Example: "Retrieve video information for swing practice")
[0382] Step 5:
[0383] Information acquisition via server
[0384] Based on the received request, the server uses the video sharing platform's API to retrieve the relevant video information. Information such as the video title and URL is obtained from the API.
[0385] Input: Request
[0386] Output: Video information (e.g., title, URL)
[0387] Step 6:
[0388] Information response
[0389] The server sends the acquired video information back to the terminal.
[0390] Input: Video information
[0391] Output: Reply message (e.g., "Video title and URL")
[0392] Step 7:
[0393] Building notifications and adjusting them to your emotions
[0394] The device adjusts the content and tone of notifications based on video information received from the server and the user's emotional state.
[0395] The device uses speech synthesis technology to generate voice messages to convey to the user. For example, if the user is anxious, it might notify them in a calm tone, "Please listen calmly. The last part of the video on the basics of swing practice explained the follow-through point. The URL is http: / / example.com."
[0396] Input: Reply message, emotional state
[0397] Output: Adjusted voice message
[0398] Step 8:
[0399] Notification to the user
[0400] The device notifies the user of the generated voice message.
[0401] Input: Pre-recorded voice message
[0402] Output: Voice notification (voice message to the user)
[0403] The above outlines the specific processing steps for implementing this invention.
[0404] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0405] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0406] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0407] [Second Embodiment]
[0408] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0409] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0410] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0411] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0412] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0413] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0414] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0415] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0416] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0417] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0418] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0419] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0420] This invention is a system for users to quickly and accurately acquire information about specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice (overly helpful advice) at practice facilities.
[0421] User voice input
[0422] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[0423] Speech-to-text conversion and analysis
[0424] The device uses speech recognition technology to convert recorded audio into text. This converted text data is used to analyze the user's intent. Natural language processing technology is used for the analysis to identify which part of which video the user wants to know about.
[0425] Request to the server
[0426] The device sends an appropriate request to the server based on the analyzed text data. This request includes information about the video or content specified by the user. For example, a request is sent to the server to re-identify a specific YouTube video and specify the "final part" of it.
[0427] Server Processing
[0428] Based on the received request, the server retrieves relevant video information from the video sharing platform. Using existing platforms such as the YouTube API, it obtains the video's title, URL, and the content of the specified portion. The server also analyzes the video's subtitles and script to identify the information the user is seeking.
[0429] Information response
[0430] The server sends the retrieved information back to the terminal. This information includes the video title, URL, and relevant content. For example, it might provide the URL of a video titled "Basic Swing Practice" and the information that "the last part explained the key points of the follow-through."
[0431] Notification to the user
[0432] The terminal notifies the user of information received from the server via voice. This information is conveyed to the user using speech synthesis technology. As a result, the user can instantly obtain information about specific techniques and methods, and easily check the instructions needed during practice. Furthermore, it enables the user to practice independently without receiving excessive advice at the practice field.
[0433] Specific example
[0434] For example, if a user says, "Show me the last part of the swing practice video I watched yesterday by a YouTuber," the device transcribes the voice command into text and sends a request to the server based on the analysis. The server retrieves information from the YouTube video, identifies the relevant section, and sends it back to the device. The device then relays this information to the user via voice, allowing the user to improve the efficiency of their practice.
[0435] The system of this invention allows users to improve the quality of their independent practice and quickly obtain necessary information. This improves practice efficiency and makes it possible to avoid excessive intervention from others.
[0436] The following describes the processing flow.
[0437] Step 1:
[0438] The user gives voice commands using an earphone microphone. For example, they input specific instructions in voice format, such as, "Show me the last part of the swing practice video I watched yesterday again."
[0439] Step 2:
[0440] The device receives and records voice input. This recorded voice data is used for subsequent processing.
[0441] Step 3:
[0442] The device uses speech recognition technology to convert recorded audio data into text data. This is done using Google's speech recognition API, among others.
[0443] Step 4:
[0444] The device analyzes what the user wants based on the converted text data. Using natural language processing technology, it extracts information to identify a specific video and its corresponding section.
[0445] Step 5:
[0446] The device sends a request to the server based on the analysis results. The request includes the URL of the video specified by the user and information about a specific part of it.
[0447] Step 6:
[0448] The server receives the request from the terminal and uses the video sharing platform's API to retrieve the relevant video information. Specifically, it retrieves the video's title, URL, and relevant content.
[0449] Step 7:
[0450] The server further analyzes the acquired video information and uses subtitles and scripts to identify what the user is looking for. For example, it might extract specific information such as "the follow-through points explained at the end."
[0451] Step 8:
[0452] The server sends the analyzed information back to the terminal. The information sent back includes the video title, URL, and a detailed description of the relevant section.
[0453] Step 9:
[0454] The terminal constructs a message to notify the user based on information received from the server. This message is then conveyed to the user in voice format using speech synthesis technology.
[0455] Step 10:
[0456] The user receives an audio notification from their device and obtains information about a specific part of the video they were looking for. Based on this information, the user can effectively proceed with their self-practice.
[0457] (Example 1)
[0458] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0459] In conventional self-practice methods, it was difficult for users to quickly and accurately obtain information on specific techniques and methods, and there was a lack of ways to avoid excessive advice, especially at the practice field. As a result, practice efficiency decreased, and it often took a long time for users to obtain the information they needed. The present invention aims to solve these problems and provide a system that allows users to improve the quality of their self-practice.
[0460] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0461] In this invention, the server includes means for recognizing voice input from a user and converting it into text data, means for analyzing the text data to determine the user's intent, means for sending a request to the server to obtain relevant information based on the user's intent, and means for notifying the user of the information received from the server by voice. This allows the user to quickly and accurately obtain the necessary information, improve the efficiency of self-practice, and minimize the impact of excessive advice.
[0462] "Voice input" refers to voice data containing instructions or information spoken by the user.
[0463] "Text data" refers to digital information obtained by converting voice input into a string format.
[0464] "Speech recognition technology" is a technology that analyzes speech information and converts its content into text data.
[0465] "Natural language processing technology" refers to techniques that analyze text data to understand the user's intentions and emotions.
[0466] A "request" is a request sent to a server to obtain specific information.
[0467] A "server" is a computer system that receives requests via a network and provides information.
[0468] A "video sharing platform" is a web service that allows users to upload, share, and watch video content.
[0469] "Speech synthesis technology" is a technology that converts text data into speech.
[0470] "Analysis results" refer to the user's intentions and necessary information obtained through the analysis of text data.
[0471] "Video information" refers to digital information about specific video content obtained from video sharing platforms.
[0472] Modes for carrying out the invention
[0473] This invention is a system for users to quickly and accurately acquire information on specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice at the practice field. Specific embodiments are described below.
[0474] User voice input
[0475] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[0476] Speech-to-text conversion
[0477] The device converts recorded audio into text using speech recognition technologies such as the Google Cloud Speech-to-Text API. This converted text data is then used to analyze the user's intent.
[0478] Text analysis
[0479] The device analyzes the converted text data using natural language processing technologies such as the Google Cloud Natural Language API. This identifies which part of which video the user wants to know about.
[0480] Sending a request to the server
[0481] The device sends an appropriate request to the server based on the analyzed text data. This request includes information about the video or content specified by the user. For example, it might include a request to specify a particular part from a particular video sharing platform.
[0482] Server Processing
[0483] Based on the received request, the server retrieves relevant video information using the YouTube API and other tools. It obtains the video title, URL, and the content of the specified section, and analyzes the subtitles and script related to that video to identify the information the user is seeking.
[0484] Information response
[0485] The server sends the retrieved information back to the terminal. This information includes the video title, URL, and relevant content. For example, it might provide the URL of a video titled "Basic Swing Practice" and the information that "the last part explained the key points of the follow-through."
[0486] Notification to the user
[0487] The device uses speech synthesis technology, such as the Google Cloud Text-to-Speech API, to notify the user of information received from the server via voice. This allows the user to instantly obtain information about specific techniques and methods, and easily check the instructions needed during practice. It also enables the user to practice independently without receiving excessive advice at the practice field.
[0488] Specific example
[0489] For example, if a user requests to see the last part of the swing practice video from the YouTuber they watched yesterday, the following process will occur:
[0490] 1. The user gives voice commands using an earphone microphone.
[0491] 2. The device records voice commands and converts them into text using speech recognition technology.
[0492] 3. The terminal parses the text and sends a request to the server.
[0493] 4. The server uses the YouTube API to retrieve the relevant video information and analyzes the content of a specific section.
[0494] 5. The server sends the retrieved information (e.g., video URL, relevant content) back to the device.
[0495] 6. The device uses speech synthesis technology to notify the user of information.
[0496] Example of a prompt
[0497] "Show me the last part of the swing practice video I saw from the YouTuber yesterday again."
[0498] This invention allows users to quickly obtain information useful for independent practice and to continue practicing efficiently.
[0499] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0500] Step 1: Recording voice input
[0501] The user transmits voice input to the device via an earphone microphone. The input voice is recorded on the device and saved as digital audio data. A specific example of this action would be the user saying, "Show me the last part of the swing practice video I watched yesterday again."
[0502] Step 2: Text conversion using speech recognition
[0503] The device sends the recorded audio data to the Google Cloud Speech-to-Text API, which converts the audio into text data. The input is the recorded audio data, and the output is the corresponding string (e.g., "Show me the last part of the swing practice video I watched yesterday again"). Specifically, this process involves analyzing the audio waveform and generating the corresponding string.
[0504] Step 3: Analyzing text data
[0505] The device sends the acquired text data to the Google Cloud Natural Language API for analysis of the user's intent. The input is the converted text data, and the output is the analysis result, including the user's intent and request. Specifically, the NLP engine analyzes the context, keywords, and intent, and recognizes it as a request regarding the final part of a particular video.
[0506] Step 4: Generate and send the request
[0507] The terminal generates and sends a request to the server to retrieve relevant information based on the analysis results. The input is the analysis results, including the user's intent, and the output is the request data sent to the server. In a specific example, a request is generated that includes "a special video URL and section for playing the final part of the swing practice."
[0508] Step 5: Collecting video information
[0509] The server processes the received request and uses the YouTube API to collect specific video information. The input is the request data sent from the device, and the output is information including the title, URL, and content of a specific part of the relevant video. Specifically, the server accesses the YouTube API to retrieve the metadata of the video and the script or subtitles of the target section.
[0510] Step 6: Return the information
[0511] The server sends the collected video information back to the device. The input is video information obtained from the YouTube API, and the output is the information sent to the device. In a specific example, data is returned that includes the URL of a video titled "Basic Swing Practice" and information about "Follow-through points explained in the last part."
[0512] Step 7: Notify the user
[0513] The device uses text-to-speech technology (Google Cloud Text-to-Speech API) to convert information received from the server into speech and notifies the user. The input is video information sent back from the server, and the output is audio data played for the user. Specifically, the device tells the user, "In the last part of the swing practice video you watched yesterday, the key points of the follow-through were explained."
[0514] Through the steps described above, the system of the present invention can quickly and accurately provide relevant video information based on the user's voice instructions. This processing flow allows the user to effectively practice independently while avoiding excessive advice.
[0515] (Application Example 1)
[0516] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0517] The operation and maintenance procedures for machinery in factories are specialized, and employees are required to acquire this information quickly and accurately. However, finding and learning the necessary information on the job is time-consuming and laborious, and can sometimes require excessive instruction. Therefore, there is a need for systems that support efficient work and self-learning among employees.
[0518] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0519] In this invention, the server includes means for recognizing voice input and converting it into text data, means for analyzing the text data to determine the user's intent, means for sending a request to an information processing device to acquire relevant information based on the user's intent, means for notifying the user of the information received from the information processing device by voice, and means for acquiring information regarding how to operate and maintain machinery in factory work. This enables employees to efficiently acquire necessary information and engage in self-learning.
[0520] "Voice input" refers to a method in which users provide information such as commands or questions using their voice via a headset microphone or similar device.
[0521] "Text data" refers to data obtained by converting voice input into characters, and is a string of characters in a format that can be analyzed by a computer.
[0522] "Analysis" is the process of analyzing text data to understand the user's intentions and requirements.
[0523] "User intent" refers to what the user is seeking through voice input, including the content of their requests and the information they desire.
[0524] A "request" is data sent to a server to retrieve relevant information based on the user's intent.
[0525] An "information processing device" refers to a system that manages and processes various types of information and data via a network.
[0526] "Notifying" refers to the process of conveying information obtained from a server or information processing device back to the user.
[0527] "Factory work" refers to all tasks performed in workplaces such as manufacturing, using machinery and equipment.
[0528] "Operating instructions" refer to the procedures and methods for making a machine or device work correctly.
[0529] "Maintenance procedures" refer to the procedures for maintenance, inspection, and repair performed to maintain the condition of machinery and equipment.
[0530] A "video sharing platform" refers to a web service that allows users to upload, play, and share videos over the internet.
[0531] This invention's system utilizes speech recognition and natural language processing technologies to provide information on how to operate and maintain machinery in a factory. Users can input voice information using an earphone microphone and quickly and accurately obtain relevant information based on that input.
[0532] System Configuration
[0533] The system consists of the following main components:
[0534] 1. Voice Input Device: The user gives voice commands using an earphone microphone. For example, they might input a voice command such as, "Show me the last part of the maintenance video I watched before again."
[0535] 2. Speech Recognition System: This system uses speech recognition technology (such as the Google Speech Recognition API) to convert recorded audio into text data. This system operates on the device.
[0536] 3. Contextual Analysis System: Natural language processing techniques (e.g., nltk or spaCy) are used to analyze text data and determine the user's intent. At this stage, relevant information is identified based on the analysis results.
[0537] 4. Server: Receives requests based on user intent and retrieves relevant video information from video sharing platforms (such as the YouTube API). It also generates a response containing the necessary information and sends it back to the device.
[0538] 5. Notification System: The terminal notifies the user of information received from the server via voice using speech synthesis technology (e.g., Pyttsx3).
[0539] Hardware and software
[0540] The following hardware and software are used to implement this system:
[0541] Hardware: Headset with microphone, smartphone or tablet, server.
[0542] Software: Google Speech Recognition API (speech recognition), nltk or spaCy (natural language processing), YouTube API (video information retrieval), Pyttsx3 (speech synthesis).
[0543] Data processing and data calculation
[0544] 1. Speech-to-text conversion: Convert user voice input into text data using the Google Speech Recognition API.
[0545] 2. Text Analysis: The generated text data is analyzed using natural language processing libraries such as nltk and spaCy to identify the user's intent.
[0546] 3. Information Request: Based on the user's request, a request for relevant video information is made to the YouTube API, and data such as the video title and URL is retrieved.
[0547] 4. Information Notification: Use Pyttsx3 to generate the acquired information in audio format and notify the user.
[0548] Specific example
[0549] If a user gives the voice command, "Show me the last part of the robot maintenance video I saw before again," the system will perform the following actions:
[0550] 1. Convert voice input to text.
[0551] 2. Analyze the converted text data to identify what information the user is looking for.
[0552] 3. Use the YouTube API to retrieve relevant video information.
[0553] 4. The acquired information (video title, URL, specified portion) is generated using speech synthesis technology and notified to the user.
[0554] Example of a prompt
[0555] User: "Show me the last part of the robot maintenance video I saw before again."
[0556] This allows users to efficiently obtain necessary information and streamline factory operations. The system reduces excessive on-site instruction and supports employees' self-directed learning.
[0557] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0558] Step 1:
[0559] The user gives voice commands using an earphone microphone. For example, they might input a voice command such as, "Show me the last part of the robot maintenance video I watched before again." The input is the user's voice data.
[0560] Step 2:
[0561] The device uses speech recognition technology (Google Speech Recognition API) to convert the user's voice data into text data. In this process, voice data is provided as input, and the converted text data is output.
[0562] Step 3:
[0563] The device uses natural language processing technology (nltk or spaCy) to analyze the generated text data and identify the user's intent. In this step, text data is the input, and the analysis outputs the user's intent (e.g., a specific part of a video to play).
[0564] Step 4:
[0565] The terminal sends a request to the server to retrieve relevant information based on the user's intent. In this step, data including the user's intent is input, and the request data to the server is output.
[0566] Step 5:
[0567] Based on the received request, the server retrieves relevant video information from the video sharing platform (YouTube API). In this step, the request data is used as input, and the video title, URL, and specific information are output.
[0568] Step 6:
[0569] The server sends the acquired information back to the terminal. In this step, video information is used as input and is sent to the terminal.
[0570] Step 7:
[0571] The terminal uses speech synthesis technology (Pyttsx3) to notify the user of the information received from the server via voice. In this step, the received video information is the input, and the output is an audio notification.
[0572] Step 8:
[0573] The user receives an audio notification from their device and reconfirms information on how to operate or maintain the relevant machine. In this step, the audio notification is provided as input, and an action is output that plays the information the user is requesting.
[0574] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0575] This invention is a system for users to quickly and accurately acquire information about specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice (overly helpful advice) at practice facilities. Furthermore, this invention can further enhance the user experience by incorporating an emotion engine that recognizes the user's emotions.
[0576] User voice input
[0577] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[0578] Speech-to-text conversion and emotion recognition
[0579] The device uses speech recognition technology to convert recorded speech into text. This utilizes a speech recognition API. The converted text data is used to analyze the user's intent. Simultaneously, an emotion engine is used to analyze the user's emotions from the speech. The emotion engine analyzes the tone and characteristics of the speech to determine the user's emotional state.
[0580] User intent analysis and request generation
[0581] The device analyzes the user's intent based on text data and sentiment data. Combining natural language processing and sentiment analysis technologies, it extracts information to identify a specific video and its associated sections. Based on the analysis results, it sends an appropriate request to the server. This request includes information about the video and content specified by the user.
[0582] Server Processing
[0583] Based on the request received from the terminal, the server uses the video sharing platform's API to retrieve relevant video information. Specifically, it obtains the video's title, URL, and relevant content. Furthermore, the server analyzes the subtitles and script related to the video to identify the information the user is seeking.
[0584] Information response and notification
[0585] The server sends the analyzed information back to the device. This information includes the video title, URL, and a detailed description of the relevant parts. Based on the information received from the server, the device constructs a message to notify the user. Using speech synthesis technology, this message is delivered to the user in voice format. The emotion engine adjusts the content and tone of the notification according to the user's emotional state. For example, if the user is anxious, the information is delivered in a calm tone; conversely, if the user is relaxed, the information is delivered in a more casual tone.
[0586] Specific example
[0587] For example, if a user says, "Show me the last part of the YouTuber's swing practice video I watched yesterday," the device transcribes the voice command into text and then uses an emotion engine to analyze the user's emotional state. Based on the analyzed text data and emotion data, the device identifies the user's intent and sends a request to the server. The server retrieves information from the YouTube video, identifies the relevant section, and sends it back to the device. The device then uses this information, along with the emotion engine's analysis results, to send a voice notification. For example, if the user is anxious, the notification might say, "Please listen calmly. The last part of the video about the basics of swing practice explained the key points of the follow-through. The URL is http: / / example.com."
[0588] The system of this invention allows users to improve the quality of their independent practice and quickly obtain necessary information. This improves practice efficiency and allows them to avoid excessive intervention from others. Furthermore, by considering the user's emotional state, it can provide a more human-like interaction.
[0589] The following describes the processing flow.
[0590] Step 1:
[0591] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again."
[0592] Step 2:
[0593] The device records voice input. This recorded voice data is used for speech recognition processing.
[0594] Step 3:
[0595] The device uses speech recognition technology to convert recorded audio data into text data. It utilizes a speech recognition API to accurately transcribe the user's speech into text.
[0596] Step 4:
[0597] The device analyzes the converted text data to identify the user's intent. Using natural language processing technology, it determines the video or specific part of the video the user is looking for.
[0598] Step 5:
[0599] The device simultaneously uses an emotion engine to analyze the user's emotional state from the voice data. It analyzes the tone of voice and speaking style to determine whether the user is excited, calm, or otherwise emotional.
[0600] Step 6:
[0601] The device sends an appropriate request to the server based on the analyzed text and sentiment data. This request includes information about the video identified by the user and the content of relevant parts of it.
[0602] Step 7:
[0603] Based on the request received from the terminal, the server uses the video sharing platform's API to retrieve the relevant video information. Specifically, it retrieves the video's title, URL, and relevant content.
[0604] Step 8:
[0605] The server further analyzes the acquired video information to identify detailed explanations of relevant parts. For example, it extracts the content explained in the final part of the video from the script or subtitles.
[0606] Step 9:
[0607] The server returns the analyzed information to the terminal. This information includes the video title, URL, and a detailed description of the relevant section.
[0608] Step 10:
[0609] The terminal constructs a message to notify the user based on information received from the server. The message is created using speech synthesis technology and adjusted according to the user's emotional state, which is analyzed by the emotion engine.
[0610] Step 11:
[0611] The device provides voice notifications and conveys acquired information to the user. For example, it might notify the user of the content and URL explained at the end of a video titled "Basic Swing Practice" in a calm or casual tone, depending on the user's emotional state.
[0612] Step 12:
[0613] Users receive audio notifications from their devices and obtain information about specific parts of the video they need. This allows users to effectively practice on their own.
[0614] (Example 2)
[0615] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0616] Conventional self-training and learning support systems have struggled to provide appropriate information quickly and accurately based on user voice input. Furthermore, they fail to consider the user's emotional state, resulting in a poor user experience. To solve these problems, the present invention aims to provide a system that efficiently transcribes user voice instructions into text, performs emotion analysis, and provides appropriate video information.
[0617] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0618] In this invention, the server includes means for recognizing voice input from a user and converting it into text data; means for analyzing the user's emotions from the text data and voice data; means for determining the user's intentions by analyzing the text data and emotion data; means for sending a request to the server to obtain relevant information based on the user's intentions; means for analyzing information obtained from the server through a video sharing platform and identifying the parts requested by the user; means for notifying the user of the information received from the server by voice; and means for adjusting the content and tone of the notification according to the user's emotional state. This enables the user to quickly and accurately obtain the necessary information and receive feedback that is appropriate to their emotions.
[0619] A "user" is an individual or organization that is using the system to obtain information.
[0620] "Voice input" refers to voice data used by a user to give instructions to a system via an earphone microphone or other voice input device.
[0621] "Text data" refers to character information converted from voice input using speech recognition technology.
[0622] "Emotional data" refers to information obtained by analyzing the user's emotional state from voice data.
[0623] "Speech recognition technology" is a technology that converts voice input into text data, and typically uses APIs or specialized software.
[0624] "Emotional analysis technology" is a technology that analyzes a user's emotional state from voice data, analyzing the tone of voice and characteristics of their speaking style.
[0625] "Natural language processing technology" is a technology that analyzes text data to understand the user's intent.
[0626] A "request" is a request for information to be retrieved, sent to a server based on the user's intent.
[0627] A "server" is a computer system that processes information based on requests and provides it to users.
[0628] A "video sharing platform" is a website or service that hosts videos that are accessible to users.
[0629] "Subtitles" are information used to display the content of a video in text format.
[0630] A "script" is text data that records the content and dialogue of a video.
[0631] "Speech synthesis technology" is a technology that converts text data into a speech format for playback.
[0632] "Notification content" refers to the specific information conveyed to the user based on information obtained from the server.
[0633] "Tone" refers to the emotion or attitude expressed in the audio of a notification.
[0634] "User intent" refers to the purpose or desired information that the user is trying to convey to the system through voice commands.
[0635] This invention is a system designed to enable users to quickly and accurately acquire information about specific techniques and methods when practicing independently. It improves the efficiency of independent learning by recognizing user voice input and providing appropriate video information. Furthermore, it can enhance the user experience by recognizing user emotions. This system includes three main elements: the user, the terminal, and the server.
[0636] Hardware and software to be used
[0637] 1. Earphone microphone (hardware): Used by the user to give voice commands.
[0638] 2. Device (hardware and software): A device for recording voice input, converting it to text, performing sentiment analysis, and intent analysis. The Google Cloud Speech-to-Text API is used for text data analysis, and the Azure Cognitive Services Emotion API is used for sentiment analysis.
[0639] 3. Server (hardware and software): Based on user instructions, it retrieves video information, analyzes that information, and identifies the necessary parts. The YouTube Data API is used for this purpose.
[0640] System operation
[0641] The user gives voice commands using an earphone microphone. For example, they might give specific instructions such as, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[0642] The device converts recorded audio into text data using the Google Cloud Speech-to-Text API. Simultaneously, it analyzes the user's emotional state from the audio data using the Azure Cognitive Services Emotion API. This text data and emotional data are then used for further analysis.
[0643] The device analyzes the user's intent using OpenAI's GPT-4 model based on text and sentiment data. Based on the analysis results, the specific video and its corresponding section that the user requests are identified. A request containing this information is then sent to the server.
[0644] The server uses the YouTube Data API to search for the specified video information and retrieves the video's title, URL, content, etc. The server then analyzes the video's subtitles and script to identify the parts requested by the user.
[0645] The server sends the analyzed information back to the device. The device uses the Google Cloud Text-to-Speech API to generate a voice message based on the information received from the server. At the same time, it adjusts the tone of the notification based on the sentiment analysis results. For example, if the user is anxious, the information will be delivered in a calm tone.
[0646] The device plays the generated voice message to the user. This allows the user to quickly and accurately obtain the necessary information.
[0647] Specific example
[0648] For example, if a user instructs, "Show me the last part of the swing practice video I watched yesterday again," the device records this audio and transcribes it using the Google Cloud Speech-to-Text API. Next, it performs sentiment analysis using the Azure Cognitive Services Emotion API. The analyzed text and sentiment data are further analyzed using the OpenAI GPT-4 model to identify a specific video segment. Based on this information, a request is sent to the server, and the relevant video information is retrieved using the YouTube Data API. The retrieved information is sent back to the device, and an audio message is generated using the Google Cloud Text-to-Speech API. The device then plays this message to the user.
[0649] For example, the user might prompt with the following message: "Show me the last part of the swing practice video I watched yesterday by a YouTuber." Based on this, the system operates and provides the user with the necessary video information.
[0650] This allows users to improve the quality of their self-practice and obtain necessary information quickly and accurately. Furthermore, feedback that takes into account the user's emotional state can provide a more human-like interaction.
[0651] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0652] Step 1:
[0653] The user gives voice commands using an earphone microphone. These voice commands are specific, such as, "Show me the last part of the swing practice video I watched yesterday again." The device records this voice and prepares for the next processing step.
[0654] Input: User's voice instructions
[0655] Output: Recorded audio data
[0656] Step 2:
[0657] The device converts the recorded audio data into text data using the Google Cloud Speech-to-Text API. The text data is the extracted content of the voice instructions as written information. The device also analyzes the user's emotional state from the audio data using the Azure Cognitive Services Emotion API.
[0658] Input: Recorded audio data
[0659] Output: Text data and sentiment data
[0660] Step 3:
[0661] The device analyzes the user's intent using OpenAI's GPT-4 model based on acquired text and sentiment data. Specifically, it identifies the particular video section the user is looking for.
[0662] Input: Text data and sentiment data
[0663] Output: Analyzed intent data (information from specific parts of the video)
[0664] Step 4:
[0665] The device sends a request to the server to retrieve appropriate information based on the analyzed intent data. The request to the server includes information about the video and content specified by the user. The request may include information about the title or specific portion of a particular video.
[0666] Input: Intent data
[0667] Output: Request to send to the server
[0668] Step 5:
[0669] Based on the request received from the terminal, the server uses the API of a video sharing platform (e.g., YouTube) to search for and retrieve the relevant video information. Specifically, it obtains the video title, URL, and relevant content, and further analyzes subtitles and scripts to identify the requested portion.
[0670] Input: Request from terminal
[0671] Output: Acquired video information (title, URL, subtitles, content of specific section)
[0672] Step 6:
[0673] The server sends the acquired video information back to the device. Based on the received information, the device generates an audio notification message using the Google Cloud Text-to-Speech API. This message includes information about the specific part of the video requested by the user. It also adjusts the tone of the notification based on the user's sentiment data.
[0674] Input: Acquired video information and emotion data
[0675] Output: Generated voice notification message
[0676] Step 7:
[0677] The device plays a generated voice notification message to the user. For example, if the user is anxious, it will notify the user in a calm tone, saying, "Please listen calmly. The last part of the video on the basics of swing practice explained the key points of the follow-through. The URL is http: / / example.com." Through voice notifications, the user can quickly and accurately obtain the information they need.
[0678] Input: Voice notification message
[0679] Output: Voice notification to the user
[0680] This allows users to improve the quality of their self-practice and quickly and accurately obtain the information they need. Furthermore, by considering the user's emotional state, a better user experience can be provided.
[0681] (Application Example 2)
[0682] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0683] While conventional learning support systems could retrieve relevant information based on user voice commands, they struggled to provide information appropriately according to the user's emotional state. This resulted in insufficient improvement in user learning efficiency and experience. Furthermore, particularly in content distribution services, there was a lack of means to quickly retrieve the learning content learners desired and to provide appropriate interactions that responded to the user's emotions.
[0684] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0685] In this invention, the server includes means for sending a request to the server to obtain relevant information based on the user's intent; means for notifying the user of the information received from the server by voice; means for determining the user's emotional state using emotion analysis technology; means for adjusting the content and tone of the notification based on the information received from the server and the user's emotional state; and means for learners to obtain learning content through voice input in a content distribution service. This enables the provision of information appropriately according to the user's emotional state, improving learning efficiency and enabling more human-like interaction.
[0686] "Voice input" refers to audio data emitted by a user using a voice collection device such as a microphone.
[0687] "Text data" refers to data obtained by converting voice input into text information using technologies such as speech recognition.
[0688] A "request" refers to the content of a request sent to a server to retrieve relevant information based on the user's intent or needs.
[0689] A "server" refers to a computer or computer system that stores and processes information over a network and provides that information to clients.
[0690] "Speech recognition technology" is a technology that analyzes speech input collected by microphones or other devices and converts it into text data.
[0691] A "video sharing platform" refers to an online service that allows users to upload, share, and watch video content over the internet.
[0692] "Emotion analysis technology" is a technology that analyzes data such as voice and text to estimate a user's emotional state.
[0693] "Notification" refers to the act of a system transmitting information or messages to a user.
[0694] "Tone" refers to the tone of voice and speaking style characteristics in messages such as notifications.
[0695] A "content distribution service" refers to a service that delivers digital content such as videos and music to users via the internet.
[0696] "Learning content" refers to educational videos, texts, and other multimedia materials that users use for learning.
[0697] The system for implementing this invention uses the following hardware and software.
[0698] Hardware and software to be used
[0699] Microphone (sound acquisition device):
[0700] It is used to collect voice input from users.
[0701] Device (smartphone, etc.):
[0702] This device performs speech recognition and emotion analysis technologies, processes data, and notifies the user of the results.
[0703] Speech recognition API:
[0704] To convert voice input into text data, you can use, for example, Google's speech recognition API.
[0705] Generative AI models:
[0706] To perform emotion analysis, for example, we can use the pipeline from the "transformers" library.
[0707] server:
[0708] The system uses the API of a video sharing platform (e.g., the YouTube Data API) to retrieve relevant video information and sends that information back to the device.
[0709] Speech synthesis technology:
[0710] This is used to notify users of information from the server as voice messages.
[0711] Explanation of the process
[0712] 1. Collection and recognition of voice input:
[0713] The user provides voice input via a microphone. For example, they might give a command like, "Show me the last part of the swing practice video I watched yesterday again."
[0714] The device collects this audio and uses speech recognition technology to convert it into text data.
[0715] 2. Emotional analysis and intention judgment:
[0716] The device uses a generative AI model to analyze the user's emotional state from text data. In this process, it uses a generative AI model (for example, the "transformers" library's pipeline) to identify emotions.
[0717] The device uses speech recognition technology to analyze text data and determine the user's intent. For example, it can understand that the user is requesting to play a specific video.
[0718] 3. Sending requests and retrieving information:
[0719] The device sends a request to the server to retrieve relevant information based on the user's intent. This request includes the search query for the required video.
[0720] The server uses an API to retrieve relevant video information (title, URL, etc.) from the video sharing platform.
[0721] 4. Information response and notification:
[0722] The server sends the acquired video information back to the terminal.
[0723] The device adjusts the content and tone of notifications based on the information it receives and the user's emotional state.
[0724] The device uses speech synthesis technology to generate an appropriate voice message and notify the user.
[0725] Specific example
[0726] For example, if a user instructs, "Show me the part of the math lesson from yesterday where we saw the proof of Fermat's Little Theorem again":
[0727] 1. User voice collection and recognition:
[0728] The user's voice commands are collected by the device and converted into text data using speech recognition technology, such as "Show me again the part of the math lesson yesterday that showed the proof of Fermat's Little Theorem."
[0729] 2. Emotional analysis and intention judgment:
[0730] Text data is analyzed using a generation AI model to identify the user's emotional state.
[0731] The device determines that the user is requesting the playback of a specific math video.
[0732] 3. Sending requests and retrieving information:
[0733] Based on user instructions, a request is sent to the server, and relevant mathematical video information is retrieved from the video sharing platform.
[0734] 4. Information response and notification:
[0735] Based on information received from the server, the device adjusts the tone of notifications according to the user's emotional state.
[0736] For example, an audio message is generated and sent to the user saying, "Here is the video proof of Fermat's Little Theorem that you requested. The URL is http: / / example.com."
[0737] Example of a prompt
[0738] Perform sentiment analysis on the AI model using the following prompt:
[0739] "Please analyze the emotion in the following sentence: 'Show me again the part of the math lesson yesterday that showed the proof of Fermat's Little Theorem.'"
[0740] This embodiment allows users to quickly obtain the necessary learning materials and receive appropriate support tailored to their emotional state.
[0741] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0742] Step 1:
[0743] Collection of user voice input
[0744] The user uses a microphone to provide voice input. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again."
[0745] Input: User's voice data
[0746] Output: Audio data is collected.
[0747] Step 2:
[0748] Text conversion of audio data
[0749] The device uses a speech recognition API to convert audio data into text data. Specifically, it uses speech recognition technology (for example, Google's speech recognition API) to convert speech into text information.
[0750] Input: Collected audio data
[0751] Output: Text data (Example: "Show me the last part of the swing practice video I watched yesterday again.")
[0752] Step 3:
[0753] Emotion analysis
[0754] The device uses a generative AI model to analyze the user's emotional state from text data. It uses emotion analysis techniques (e.g., the transformers library's pipeline) to analyze the tone and characteristics of speech to determine the emotional state.
[0755] Input: Text data
[0756] Output: Emotion analysis results (e.g., "anxiety")
[0757] Step 4:
[0758] Intent analysis and request generation
[0759] The device analyzes the user's intent based on text data and sentiment analysis results. Specifically, it uses natural language processing technology to identify what the user is looking for.
[0760] Based on the analysis results, the terminal sends a request to the server to retrieve relevant information.
[0761] Input: Text data, sentiment analysis results
[0762] Output: Request (Example: "Retrieve video information for swing practice")
[0763] Step 5:
[0764] Information acquisition via server
[0765] Based on the received request, the server uses the video sharing platform's API to retrieve the relevant video information. Information such as the video title and URL is obtained from the API.
[0766] Input: Request
[0767] Output: Video information (e.g., title, URL)
[0768] Step 6:
[0769] Information response
[0770] The server sends the acquired video information back to the terminal.
[0771] Input: Video information
[0772] Output: Reply message (e.g., "Video title and URL")
[0773] Step 7:
[0774] Building notifications and adjusting them to your emotions
[0775] The device adjusts the content and tone of notifications based on video information received from the server and the user's emotional state.
[0776] The device uses speech synthesis technology to generate voice messages to convey to the user. For example, if the user is anxious, it might notify them in a calm tone, "Please listen calmly. The last part of the video on the basics of swing practice explained the follow-through point. The URL is http: / / example.com."
[0777] Input: Reply message, emotional state
[0778] Output: Adjusted voice message
[0779] Step 8:
[0780] Notification to the user
[0781] The device notifies the user of the generated voice message.
[0782] Input: Pre-recorded voice message
[0783] Output: Voice notification (voice message to the user)
[0784] The above outlines the specific processing steps for implementing this invention.
[0785] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0786] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0787] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0788] [Third Embodiment]
[0789] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0790] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0791] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0792] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0793] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0794] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0795] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0796] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0797] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0798] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0799] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0800] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0801] This invention is a system for users to quickly and accurately acquire information about specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice (overly helpful advice) at practice facilities.
[0802] User voice input
[0803] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[0804] Speech-to-text conversion and analysis
[0805] The device uses speech recognition technology to convert recorded audio into text. This converted text data is used to analyze the user's intent. Natural language processing technology is used for the analysis to identify which part of which video the user wants to know about.
[0806] Request to the server
[0807] The device sends an appropriate request to the server based on the analyzed text data. This request includes information about the video or content specified by the user. For example, a request is sent to the server to re-identify a specific YouTube video and specify the "final part" of it.
[0808] Server Processing
[0809] Based on the received request, the server retrieves relevant video information from the video sharing platform. Using existing platforms such as the YouTube API, it obtains the video's title, URL, and the content of the specified portion. The server also analyzes the video's subtitles and script to identify the information the user is seeking.
[0810] Information response
[0811] The server sends the retrieved information back to the terminal. This information includes the video title, URL, and relevant content. For example, it might provide the URL of a video titled "Basic Swing Practice" and the information that "the last part explained the key points of the follow-through."
[0812] Notification to the user
[0813] The terminal notifies the user of information received from the server via voice. This information is conveyed to the user using speech synthesis technology. As a result, the user can instantly obtain information about specific techniques and methods, and easily check the instructions needed during practice. Furthermore, it enables the user to practice independently without receiving excessive advice at the practice field.
[0814] Specific example
[0815] For example, if a user says, "Show me the last part of the swing practice video I watched yesterday by a YouTuber," the device transcribes the voice command into text and sends a request to the server based on the analysis. The server retrieves information from the YouTube video, identifies the relevant section, and sends it back to the device. The device then relays this information to the user via voice, allowing the user to improve the efficiency of their practice.
[0816] The system of this invention allows users to improve the quality of their independent practice and quickly obtain necessary information. This improves practice efficiency and makes it possible to avoid excessive intervention from others.
[0817] The following describes the processing flow.
[0818] Step 1:
[0819] The user gives voice commands using an earphone microphone. For example, they input specific instructions in voice format, such as, "Show me the last part of the swing practice video I watched yesterday again."
[0820] Step 2:
[0821] The device receives and records voice input. This recorded voice data is used for subsequent processing.
[0822] Step 3:
[0823] The device uses speech recognition technology to convert recorded audio data into text data. This is done using Google's speech recognition API, among others.
[0824] Step 4:
[0825] The device analyzes what the user wants based on the converted text data. Using natural language processing technology, it extracts information to identify a specific video and its corresponding section.
[0826] Step 5:
[0827] The device sends a request to the server based on the analysis results. The request includes the URL of the video specified by the user and information about a specific part of it.
[0828] Step 6:
[0829] The server receives the request from the terminal and uses the video sharing platform's API to retrieve the relevant video information. Specifically, it retrieves the video's title, URL, and relevant content.
[0830] Step 7:
[0831] The server further analyzes the acquired video information and uses subtitles and scripts to identify what the user is looking for. For example, it might extract specific information such as "the follow-through points explained at the end."
[0832] Step 8:
[0833] The server sends the analyzed information back to the terminal. The information sent back includes the video title, URL, and a detailed description of the relevant section.
[0834] Step 9:
[0835] The terminal constructs a message to notify the user based on information received from the server. This message is then conveyed to the user in voice format using speech synthesis technology.
[0836] Step 10:
[0837] The user receives an audio notification from their device and obtains information about a specific part of the video they were looking for. Based on this information, the user can effectively proceed with their self-practice.
[0838] (Example 1)
[0839] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0840] In conventional self-practice methods, it was difficult for users to quickly and accurately obtain information on specific techniques and methods, and there was a lack of ways to avoid excessive advice, especially at the practice field. As a result, practice efficiency decreased, and it often took a long time for users to obtain the information they needed. The present invention aims to solve these problems and provide a system that allows users to improve the quality of their self-practice.
[0841] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0842] In this invention, the server includes means for recognizing voice input from a user and converting it into text data, means for analyzing the text data to determine the user's intent, means for sending a request to the server to obtain relevant information based on the user's intent, and means for notifying the user of the information received from the server by voice. This allows the user to quickly and accurately obtain the necessary information, improve the efficiency of self-practice, and minimize the impact of excessive advice.
[0843] "Voice input" refers to voice data containing instructions or information spoken by the user.
[0844] "Text data" refers to digital information obtained by converting voice input into a string format.
[0845] "Speech recognition technology" is a technology that analyzes speech information and converts its content into text data.
[0846] "Natural language processing technology" refers to techniques that analyze text data to understand the user's intentions and emotions.
[0847] A "request" is a request sent to a server to obtain specific information.
[0848] A "server" is a computer system that receives requests via a network and provides information.
[0849] A "video sharing platform" is a web service that allows users to upload, share, and watch video content.
[0850] "Speech synthesis technology" is a technology that converts text data into speech.
[0851] "Analysis results" refer to the user's intentions and necessary information obtained through the analysis of text data.
[0852] "Video information" refers to digital information about specific video content obtained from video sharing platforms.
[0853] Modes for carrying out the invention
[0854] This invention is a system for users to quickly and accurately acquire information on specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice at the practice field. Specific embodiments are described below.
[0855] User voice input
[0856] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[0857] Speech-to-text conversion
[0858] The device converts recorded audio into text using speech recognition technologies such as the Google Cloud Speech-to-Text API. This converted text data is then used to analyze the user's intent.
[0859] Text analysis
[0860] The device analyzes the converted text data using natural language processing technologies such as the Google Cloud Natural Language API. This identifies which part of which video the user wants to know about.
[0861] Sending a request to the server
[0862] The device sends an appropriate request to the server based on the analyzed text data. This request includes information about the video or content specified by the user. For example, it might include a request to specify a particular part from a particular video sharing platform.
[0863] Server Processing
[0864] Based on the received request, the server retrieves relevant video information using the YouTube API and other tools. It obtains the video title, URL, and the content of the specified section, and analyzes the subtitles and script related to that video to identify the information the user is seeking.
[0865] Information response
[0866] The server sends the retrieved information back to the terminal. This information includes the video title, URL, and relevant content. For example, it might provide the URL of a video titled "Basic Swing Practice" and the information that "the last part explained the key points of the follow-through."
[0867] Notification to the user
[0868] The device uses speech synthesis technology, such as the Google Cloud Text-to-Speech API, to notify the user of information received from the server via voice. This allows the user to instantly obtain information about specific techniques and methods, and easily check the instructions needed during practice. It also enables the user to practice independently without receiving excessive advice at the practice field.
[0869] Specific example
[0870] For example, if a user requests to see the last part of the swing practice video from the YouTuber they watched yesterday, the following process will occur:
[0871] 1. The user gives voice commands using an earphone microphone.
[0872] 2. The device records voice commands and converts them into text using speech recognition technology.
[0873] 3. The terminal parses the text and sends a request to the server.
[0874] 4. The server uses the YouTube API to retrieve the relevant video information and analyzes the content of a specific section.
[0875] 5. The server sends the retrieved information (e.g., video URL, relevant content) back to the device.
[0876] 6. The device uses speech synthesis technology to notify the user of information.
[0877] Example of a prompt
[0878] "Show me the last part of the swing practice video I saw from the YouTuber yesterday again."
[0879] This invention allows users to quickly obtain information useful for independent practice and to continue practicing efficiently.
[0880] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0881] Step 1: Recording voice input
[0882] The user transmits voice input to the device via an earphone microphone. The input voice is recorded on the device and saved as digital audio data. A specific example of this action would be the user saying, "Show me the last part of the swing practice video I watched yesterday again."
[0883] Step 2: Text conversion using speech recognition
[0884] The device sends the recorded audio data to the Google Cloud Speech-to-Text API, which converts the audio into text data. The input is the recorded audio data, and the output is the corresponding string (e.g., "Show me the last part of the swing practice video I watched yesterday again"). Specifically, this process involves analyzing the audio waveform and generating the corresponding string.
[0885] Step 3: Analyzing text data
[0886] The device sends the acquired text data to the Google Cloud Natural Language API for analysis of the user's intent. The input is the converted text data, and the output is the analysis result, including the user's intent and request. Specifically, the NLP engine analyzes the context, keywords, and intent, and recognizes it as a request regarding the final part of a particular video.
[0887] Step 4: Generate and send the request
[0888] The terminal generates and sends a request to the server to retrieve relevant information based on the analysis results. The input is the analysis results, including the user's intent, and the output is the request data sent to the server. In a specific example, a request is generated that includes "a special video URL and section for playing the final part of the swing practice."
[0889] Step 5: Collecting video information
[0890] The server processes the received request and uses the YouTube API to collect specific video information. The input is the request data sent from the device, and the output is information including the title, URL, and content of a specific part of the relevant video. Specifically, the server accesses the YouTube API to retrieve the metadata of the video and the script or subtitles of the target section.
[0891] Step 6: Return the information
[0892] The server sends the collected video information back to the device. The input is video information obtained from the YouTube API, and the output is the information sent to the device. In a specific example, data is returned that includes the URL of a video titled "Basic Swing Practice" and information about "Follow-through points explained in the last part."
[0893] Step 7: Notify the user
[0894] The device uses text-to-speech technology (Google Cloud Text-to-Speech API) to convert information received from the server into speech and notifies the user. The input is video information sent back from the server, and the output is audio data played for the user. Specifically, the device tells the user, "In the last part of the swing practice video you watched yesterday, the key points of the follow-through were explained."
[0895] Through the steps described above, the system of the present invention can quickly and accurately provide relevant video information based on the user's voice instructions. This processing flow allows the user to effectively practice independently while avoiding excessive advice.
[0896] (Application Example 1)
[0897] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0898] The operation and maintenance procedures for machinery in factories are specialized, and employees are required to acquire this information quickly and accurately. However, finding and learning the necessary information on the job is time-consuming and laborious, and can sometimes require excessive instruction. Therefore, there is a need for systems that support efficient work and self-learning among employees.
[0899] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0900] In this invention, the server includes means for recognizing voice input and converting it into text data, means for analyzing the text data to determine the user's intent, means for sending a request to an information processing device to acquire relevant information based on the user's intent, means for notifying the user of the information received from the information processing device by voice, and means for acquiring information regarding how to operate and maintain machinery in factory work. This enables employees to efficiently acquire necessary information and engage in self-learning.
[0901] "Voice input" refers to a method in which users provide information such as commands or questions using their voice via a headset microphone or similar device.
[0902] "Text data" refers to data obtained by converting voice input into characters, and is a string of characters in a format that can be analyzed by a computer.
[0903] "Analysis" is the process of analyzing text data to understand the user's intentions and requirements.
[0904] "User intent" refers to what the user is seeking through voice input, including the content of their requests and the information they desire.
[0905] A "request" is data sent to a server to retrieve relevant information based on the user's intent.
[0906] An "information processing device" refers to a system that manages and processes various types of information and data via a network.
[0907] "Notifying" refers to the process of conveying information obtained from a server or information processing device back to the user.
[0908] "Factory work" refers to all tasks performed in workplaces such as manufacturing, using machinery and equipment.
[0909] "Operating instructions" refer to the procedures and methods for making a machine or device work correctly.
[0910] "Maintenance procedures" refer to the procedures for maintenance, inspection, and repair performed to maintain the condition of machinery and equipment.
[0911] A "video sharing platform" refers to a web service that allows users to upload, play, and share videos over the internet.
[0912] This invention's system utilizes speech recognition and natural language processing technologies to provide information on how to operate and maintain machinery in a factory. Users can input voice information using an earphone microphone and quickly and accurately obtain relevant information based on that input.
[0913] System Configuration
[0914] The system consists of the following main components:
[0915] 1. Voice Input Device: The user gives voice commands using an earphone microphone. For example, they might input a voice command such as, "Show me the last part of the maintenance video I watched before again."
[0916] 2. Speech Recognition System: This system uses speech recognition technology (such as the Google Speech Recognition API) to convert recorded audio into text data. This system operates on the device.
[0917] 3. Contextual Analysis System: Natural language processing techniques (e.g., nltk or spaCy) are used to analyze text data and determine the user's intent. At this stage, relevant information is identified based on the analysis results.
[0918] 4. Server: Receives requests based on user intent and retrieves relevant video information from video sharing platforms (such as the YouTube API). It also generates a response containing the necessary information and sends it back to the device.
[0919] 5. Notification System: The terminal notifies the user of information received from the server via voice using speech synthesis technology (e.g., Pyttsx3).
[0920] Hardware and software
[0921] The following hardware and software are used to implement this system:
[0922] Hardware: Headset with microphone, smartphone or tablet, server.
[0923] Software: Google Speech Recognition API (speech recognition), nltk or spaCy (natural language processing), YouTube API (video information retrieval), Pyttsx3 (speech synthesis).
[0924] Data processing and data calculation
[0925] 1. Speech-to-text conversion: Convert user voice input into text data using the Google Speech Recognition API.
[0926] 2. Text Analysis: The generated text data is analyzed using natural language processing libraries such as nltk and spaCy to identify the user's intent.
[0927] 3. Information Request: Based on the user's request, a request for relevant video information is made to the YouTube API, and data such as the video title and URL is retrieved.
[0928] 4. Information Notification: Use Pyttsx3 to generate the acquired information in audio format and notify the user.
[0929] Specific example
[0930] If a user gives the voice command, "Show me the last part of the robot maintenance video I saw before again," the system will perform the following actions:
[0931] 1. Convert voice input to text.
[0932] 2. Analyze the converted text data to identify what information the user is looking for.
[0933] 3. Use the YouTube API to retrieve relevant video information.
[0934] 4. The acquired information (video title, URL, specified portion) is generated using speech synthesis technology and notified to the user.
[0935] Example of a prompt
[0936] User: "Show me the last part of the robot maintenance video I saw before again."
[0937] This allows users to efficiently obtain necessary information and streamline factory operations. The system reduces excessive on-site instruction and supports employees' self-directed learning.
[0938] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0939] Step 1:
[0940] The user gives voice commands using an earphone microphone. For example, they might input a voice command such as, "Show me the last part of the robot maintenance video I watched before again." The input is the user's voice data.
[0941] Step 2:
[0942] The device uses speech recognition technology (Google Speech Recognition API) to convert the user's voice data into text data. In this process, voice data is provided as input, and the converted text data is output.
[0943] Step 3:
[0944] The device uses natural language processing technology (nltk or spaCy) to analyze the generated text data and identify the user's intent. In this step, text data is the input, and the analysis outputs the user's intent (e.g., a specific part of a video to play).
[0945] Step 4:
[0946] The terminal sends a request to the server to retrieve relevant information based on the user's intent. In this step, data including the user's intent is input, and the request data to the server is output.
[0947] Step 5:
[0948] Based on the received request, the server retrieves relevant video information from the video sharing platform (YouTube API). In this step, the request data is used as input, and the video title, URL, and specific information are output.
[0949] Step 6:
[0950] The server sends the acquired information back to the terminal. In this step, video information is used as input and is sent to the terminal.
[0951] Step 7:
[0952] The terminal uses speech synthesis technology (Pyttsx3) to notify the user of the information received from the server via voice. In this step, the received video information is the input, and the output is an audio notification.
[0953] Step 8:
[0954] The user receives an audio notification from their device and reconfirms information on how to operate or maintain the relevant machine. In this step, the audio notification is provided as input, and an action is output that plays the information the user is requesting.
[0955] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0956] This invention is a system for users to quickly and accurately acquire information about specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice (overly helpful advice) at practice facilities. Furthermore, this invention can further enhance the user experience by incorporating an emotion engine that recognizes the user's emotions.
[0957] User voice input
[0958] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[0959] Speech-to-text conversion and emotion recognition
[0960] The device uses speech recognition technology to convert recorded speech into text. This utilizes a speech recognition API. The converted text data is used to analyze the user's intent. Simultaneously, an emotion engine is used to analyze the user's emotions from the speech. The emotion engine analyzes the tone and characteristics of the speech to determine the user's emotional state.
[0961] User intent analysis and request generation
[0962] The device analyzes the user's intent based on text data and sentiment data. Combining natural language processing and sentiment analysis technologies, it extracts information to identify a specific video and its associated sections. Based on the analysis results, it sends an appropriate request to the server. This request includes information about the video and content specified by the user.
[0963] Server Processing
[0964] Based on the request received from the terminal, the server uses the video sharing platform's API to retrieve relevant video information. Specifically, it obtains the video's title, URL, and relevant content. Furthermore, the server analyzes the subtitles and script related to the video to identify the information the user is seeking.
[0965] Information response and notification
[0966] The server sends the analyzed information back to the device. This information includes the video title, URL, and a detailed description of the relevant parts. Based on the information received from the server, the device constructs a message to notify the user. Using speech synthesis technology, this message is delivered to the user in voice format. The emotion engine adjusts the content and tone of the notification according to the user's emotional state. For example, if the user is anxious, the information is delivered in a calm tone; conversely, if the user is relaxed, the information is delivered in a more casual tone.
[0967] Specific example
[0968] For example, if a user says, "Show me the last part of the YouTuber's swing practice video I watched yesterday," the device transcribes the voice command into text and then uses an emotion engine to analyze the user's emotional state. Based on the analyzed text data and emotion data, the device identifies the user's intent and sends a request to the server. The server retrieves information from the YouTube video, identifies the relevant section, and sends it back to the device. The device then uses this information, along with the emotion engine's analysis results, to send a voice notification. For example, if the user is anxious, the notification might say, "Please listen calmly. The last part of the video about the basics of swing practice explained the key points of the follow-through. The URL is http: / / example.com."
[0969] The system of this invention allows users to improve the quality of their independent practice and quickly obtain necessary information. This improves practice efficiency and allows them to avoid excessive intervention from others. Furthermore, by considering the user's emotional state, it can provide a more human-like interaction.
[0970] The following describes the processing flow.
[0971] Step 1:
[0972] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again."
[0973] Step 2:
[0974] The device records voice input. This recorded voice data is used for speech recognition processing.
[0975] Step 3:
[0976] The device uses speech recognition technology to convert recorded audio data into text data. It utilizes a speech recognition API to accurately transcribe the user's speech into text.
[0977] Step 4:
[0978] The device analyzes the converted text data to identify the user's intent. Using natural language processing technology, it determines the video or specific part of the video the user is looking for.
[0979] Step 5:
[0980] The device simultaneously uses an emotion engine to analyze the user's emotional state from the voice data. It analyzes the tone of voice and speaking style to determine whether the user is excited, calm, or otherwise emotional.
[0981] Step 6:
[0982] The device sends an appropriate request to the server based on the analyzed text and sentiment data. This request includes information about the video identified by the user and the content of relevant parts of it.
[0983] Step 7:
[0984] Based on the request received from the terminal, the server uses the video sharing platform's API to retrieve the relevant video information. Specifically, it retrieves the video's title, URL, and relevant content.
[0985] Step 8:
[0986] The server further analyzes the acquired video information to identify detailed explanations of relevant parts. For example, it extracts the content explained in the final part of the video from the script or subtitles.
[0987] Step 9:
[0988] The server returns the analyzed information to the terminal. This information includes the video title, URL, and a detailed description of the relevant section.
[0989] Step 10:
[0990] The terminal constructs a message to notify the user based on information received from the server. The message is created using speech synthesis technology and adjusted according to the user's emotional state, which is analyzed by the emotion engine.
[0991] Step 11:
[0992] The device provides voice notifications and conveys acquired information to the user. For example, it might notify the user of the content and URL explained at the end of a video titled "Basic Swing Practice" in a calm or casual tone, depending on the user's emotional state.
[0993] Step 12:
[0994] Users receive audio notifications from their devices and obtain information about specific parts of the video they need. This allows users to effectively practice on their own.
[0995] (Example 2)
[0996] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0997] Conventional self-training and learning support systems have struggled to provide appropriate information quickly and accurately based on user voice input. Furthermore, they fail to consider the user's emotional state, resulting in a poor user experience. To solve these problems, the present invention aims to provide a system that efficiently transcribes user voice instructions into text, performs emotion analysis, and provides appropriate video information.
[0998] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0999] In this invention, the server includes means for recognizing voice input from a user and converting it into text data; means for analyzing the user's emotions from the text data and voice data; means for determining the user's intentions by analyzing the text data and emotion data; means for sending a request to the server to obtain relevant information based on the user's intentions; means for analyzing information obtained from the server through a video sharing platform and identifying the parts requested by the user; means for notifying the user of the information received from the server by voice; and means for adjusting the content and tone of the notification according to the user's emotional state. This enables the user to quickly and accurately obtain the necessary information and receive feedback that is appropriate to their emotions.
[1000] A "user" is an individual or organization that is using the system to obtain information.
[1001] "Voice input" refers to voice data used by a user to give instructions to a system via an earphone microphone or other voice input device.
[1002] "Text data" refers to character information converted from voice input using speech recognition technology.
[1003] "Emotional data" refers to information obtained by analyzing the user's emotional state from voice data.
[1004] "Speech recognition technology" is a technology that converts voice input into text data, and typically uses APIs or specialized software.
[1005] "Emotional analysis technology" is a technology that analyzes a user's emotional state from voice data, analyzing the tone of voice and characteristics of their speaking style.
[1006] "Natural language processing technology" is a technology that analyzes text data to understand the user's intent.
[1007] A "request" is a request for information to be retrieved, sent to a server based on the user's intent.
[1008] A "server" is a computer system that processes information based on requests and provides it to users.
[1009] A "video sharing platform" is a website or service that hosts videos that are accessible to users.
[1010] "Subtitles" are information used to display the content of a video in text format.
[1011] A "script" is text data that records the content and dialogue of a video.
[1012] "Speech synthesis technology" is a technology that converts text data into a speech format for playback.
[1013] "Notification content" refers to the specific information conveyed to the user based on information obtained from the server.
[1014] "Tone" refers to the emotion or attitude expressed in the audio of a notification.
[1015] "User intent" refers to the purpose or desired information that the user is trying to convey to the system through voice commands.
[1016] This invention is a system designed to enable users to quickly and accurately acquire information about specific techniques and methods when practicing independently. It improves the efficiency of independent learning by recognizing user voice input and providing appropriate video information. Furthermore, it can enhance the user experience by recognizing user emotions. This system includes three main elements: the user, the terminal, and the server.
[1017] Hardware and software to be used
[1018] 1. Earphone microphone (hardware): Used by the user to give voice commands.
[1019] 2. Device (hardware and software): A device for recording voice input, converting it to text, performing sentiment analysis, and intent analysis. The Google Cloud Speech-to-Text API is used for text data analysis, and the Azure Cognitive Services Emotion API is used for sentiment analysis.
[1020] 3. Server (hardware and software): Based on user instructions, it retrieves video information, analyzes that information, and identifies the necessary parts. The YouTube Data API is used for this purpose.
[1021] System operation
[1022] The user gives voice commands using an earphone microphone. For example, they might give specific instructions such as, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[1023] The device converts recorded audio into text data using the Google Cloud Speech-to-Text API. Simultaneously, it analyzes the user's emotional state from the audio data using the Azure Cognitive Services Emotion API. This text data and emotional data are then used for further analysis.
[1024] The device analyzes the user's intent using OpenAI's GPT-4 model based on text and sentiment data. Based on the analysis results, the specific video and its corresponding section that the user requests are identified. A request containing this information is then sent to the server.
[1025] The server uses the YouTube Data API to search for the specified video information and retrieves the video's title, URL, content, etc. The server then analyzes the video's subtitles and script to identify the parts requested by the user.
[1026] The server sends the analyzed information back to the device. The device uses the Google Cloud Text-to-Speech API to generate a voice message based on the information received from the server. At the same time, it adjusts the tone of the notification based on the sentiment analysis results. For example, if the user is anxious, the information will be delivered in a calm tone.
[1027] The device plays the generated voice message to the user. This allows the user to quickly and accurately obtain the necessary information.
[1028] Specific example
[1029] For example, if a user instructs, "Show me the last part of the swing practice video I watched yesterday again," the device records this audio and transcribes it using the Google Cloud Speech-to-Text API. Next, it performs sentiment analysis using the Azure Cognitive Services Emotion API. The analyzed text and sentiment data are further analyzed using the OpenAI GPT-4 model to identify a specific video segment. Based on this information, a request is sent to the server, and the relevant video information is retrieved using the YouTube Data API. The retrieved information is sent back to the device, and an audio message is generated using the Google Cloud Text-to-Speech API. The device then plays this message to the user.
[1030] For example, the user might prompt with the following message: "Show me the last part of the swing practice video I watched yesterday by a YouTuber." Based on this, the system operates and provides the user with the necessary video information.
[1031] This allows users to improve the quality of their self-practice and obtain necessary information quickly and accurately. Furthermore, feedback that takes into account the user's emotional state can provide a more human-like interaction.
[1032] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1033] Step 1:
[1034] The user gives voice commands using an earphone microphone. These voice commands are specific, such as, "Show me the last part of the swing practice video I watched yesterday again." The device records this voice and prepares for the next processing step.
[1035] Input: User's voice instructions
[1036] Output: Recorded audio data
[1037] Step 2:
[1038] The device converts the recorded audio data into text data using the Google Cloud Speech-to-Text API. The text data is the extracted content of the voice instructions as written information. The device also analyzes the user's emotional state from the audio data using the Azure Cognitive Services Emotion API.
[1039] Input: Recorded audio data
[1040] Output: Text data and sentiment data
[1041] Step 3:
[1042] The device analyzes the user's intent using OpenAI's GPT-4 model based on acquired text and sentiment data. Specifically, it identifies the particular video section the user is looking for.
[1043] Input: Text data and sentiment data
[1044] Output: Analyzed intent data (information from specific parts of the video)
[1045] Step 4:
[1046] The device sends a request to the server to retrieve appropriate information based on the analyzed intent data. The request to the server includes information about the video and content specified by the user. The request may include information about the title or specific portion of a particular video.
[1047] Input: Intent data
[1048] Output: Request to send to the server
[1049] Step 5:
[1050] Based on the request received from the terminal, the server uses the API of a video sharing platform (e.g., YouTube) to search for and retrieve the relevant video information. Specifically, it obtains the video title, URL, and relevant content, and further analyzes subtitles and scripts to identify the requested portion.
[1051] Input: Request from terminal
[1052] Output: Acquired video information (title, URL, subtitles, content of specific section)
[1053] Step 6:
[1054] The server sends the acquired video information back to the device. Based on the received information, the device generates an audio notification message using the Google Cloud Text-to-Speech API. This message includes information about the specific part of the video requested by the user. It also adjusts the tone of the notification based on the user's sentiment data.
[1055] Input: Acquired video information and emotion data
[1056] Output: Generated voice notification message
[1057] Step 7:
[1058] The device plays a generated voice notification message to the user. For example, if the user is anxious, it will notify the user in a calm tone, saying, "Please listen calmly. The last part of the video on the basics of swing practice explained the key points of the follow-through. The URL is http: / / example.com." Through voice notifications, the user can quickly and accurately obtain the information they need.
[1059] Input: Voice notification message
[1060] Output: Voice notification to the user
[1061] This allows users to improve the quality of their self-practice and quickly and accurately obtain the information they need. Furthermore, by considering the user's emotional state, a better user experience can be provided.
[1062] (Application Example 2)
[1063] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1064] While conventional learning support systems could retrieve relevant information based on user voice commands, they struggled to provide information appropriately according to the user's emotional state. This resulted in insufficient improvement in user learning efficiency and experience. Furthermore, particularly in content distribution services, there was a lack of means to quickly retrieve the learning content learners desired and to provide appropriate interactions that responded to the user's emotions.
[1065] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1066] In this invention, the server includes means for sending a request to the server to obtain relevant information based on the user's intent; means for notifying the user of the information received from the server by voice; means for determining the user's emotional state using emotion analysis technology; means for adjusting the content and tone of the notification based on the information received from the server and the user's emotional state; and means for learners to obtain learning content through voice input in a content distribution service. This enables the provision of information appropriately according to the user's emotional state, improving learning efficiency and enabling more human-like interaction.
[1067] "Voice input" refers to audio data emitted by a user using a voice collection device such as a microphone.
[1068] "Text data" refers to data obtained by converting voice input into text information using technologies such as speech recognition.
[1069] A "request" refers to the content of a request sent to a server to retrieve relevant information based on the user's intent or needs.
[1070] A "server" refers to a computer or computer system that stores and processes information over a network and provides that information to clients.
[1071] "Speech recognition technology" is a technology that analyzes speech input collected by microphones or other devices and converts it into text data.
[1072] A "video sharing platform" refers to an online service that allows users to upload, share, and watch video content over the internet.
[1073] "Emotion analysis technology" is a technology that analyzes data such as voice and text to estimate a user's emotional state.
[1074] "Notification" refers to the act of a system transmitting information or messages to a user.
[1075] "Tone" refers to the tone of voice and speaking style characteristics in messages such as notifications.
[1076] A "content distribution service" refers to a service that delivers digital content such as videos and music to users via the internet.
[1077] "Learning content" refers to educational videos, texts, and other multimedia materials that users use for learning.
[1078] The system for implementing this invention uses the following hardware and software.
[1079] Hardware and software to be used
[1080] Microphone (sound acquisition device):
[1081] It is used to collect voice input from users.
[1082] Device (smartphone, etc.):
[1083] This device performs speech recognition and emotion analysis technologies, processes data, and notifies the user of the results.
[1084] Speech recognition API:
[1085] To convert voice input into text data, you can use, for example, Google's speech recognition API.
[1086] Generative AI models:
[1087] To perform emotion analysis, for example, we can use the pipeline from the "transformers" library.
[1088] server:
[1089] The system uses the API of a video sharing platform (e.g., the YouTube Data API) to retrieve relevant video information and sends that information back to the device.
[1090] Speech synthesis technology:
[1091] This is used to notify users of information from the server as voice messages.
[1092] Explanation of the process
[1093] 1. Collection and recognition of voice input:
[1094] The user provides voice input via a microphone. For example, they might give a command like, "Show me the last part of the swing practice video I watched yesterday again."
[1095] The device collects this audio and uses speech recognition technology to convert it into text data.
[1096] 2. Emotional analysis and intention judgment:
[1097] The device uses a generative AI model to analyze the user's emotional state from text data. In this process, it uses a generative AI model (for example, the "transformers" library's pipeline) to identify emotions.
[1098] The device uses speech recognition technology to analyze text data and determine the user's intent. For example, it can understand that the user is requesting to play a specific video.
[1099] 3. Sending requests and retrieving information:
[1100] The device sends a request to the server to retrieve relevant information based on the user's intent. This request includes the search query for the required video.
[1101] The server uses an API to retrieve relevant video information (title, URL, etc.) from the video sharing platform.
[1102] 4. Information response and notification:
[1103] The server sends the acquired video information back to the terminal.
[1104] The device adjusts the content and tone of notifications based on the information it receives and the user's emotional state.
[1105] The device uses speech synthesis technology to generate an appropriate voice message and notify the user.
[1106] Specific example
[1107] For example, if a user instructs, "Show me the part of the math lesson from yesterday where we saw the proof of Fermat's Little Theorem again":
[1108] 1. User voice collection and recognition:
[1109] The user's voice commands are collected by the device and converted into text data using speech recognition technology, such as "Show me again the part of the math lesson yesterday that showed the proof of Fermat's Little Theorem."
[1110] 2. Emotional analysis and intention judgment:
[1111] Text data is analyzed using a generation AI model to identify the user's emotional state.
[1112] The device determines that the user is requesting the playback of a specific math video.
[1113] 3. Sending requests and retrieving information:
[1114] Based on user instructions, a request is sent to the server, and relevant mathematical video information is retrieved from the video sharing platform.
[1115] 4. Information response and notification:
[1116] Based on information received from the server, the device adjusts the tone of notifications according to the user's emotional state.
[1117] For example, an audio message is generated and sent to the user saying, "Here is the video proof of Fermat's Little Theorem that you requested. The URL is http: / / example.com."
[1118] Example of a prompt
[1119] Perform sentiment analysis on the AI model using the following prompt:
[1120] "Please analyze the emotion in the following sentence: 'Show me again the part of the math lesson yesterday that showed the proof of Fermat's Little Theorem.'"
[1121] This embodiment allows users to quickly obtain the necessary learning materials and receive appropriate support tailored to their emotional state.
[1122] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1123] Step 1:
[1124] Collection of user voice input
[1125] The user uses a microphone to provide voice input. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again."
[1126] Input: User's voice data
[1127] Output: Audio data is collected.
[1128] Step 2:
[1129] Text conversion of audio data
[1130] The device uses a speech recognition API to convert audio data into text data. Specifically, it uses speech recognition technology (for example, Google's speech recognition API) to convert speech into text information.
[1131] Input: Collected audio data
[1132] Output: Text data (Example: "Show me the last part of the swing practice video I watched yesterday again.")
[1133] Step 3:
[1134] Emotion analysis
[1135] The device uses a generative AI model to analyze the user's emotional state from text data. It uses emotion analysis techniques (e.g., the transformers library's pipeline) to analyze the tone and characteristics of speech to determine the emotional state.
[1136] Input: Text data
[1137] Output: Emotion analysis results (e.g., "anxiety")
[1138] Step 4:
[1139] Intent analysis and request generation
[1140] The device analyzes the user's intent based on text data and sentiment analysis results. Specifically, it uses natural language processing technology to identify what the user is looking for.
[1141] Based on the analysis results, the terminal sends a request to the server to retrieve relevant information.
[1142] Input: Text data, sentiment analysis results
[1143] Output: Request (Example: "Retrieve video information for swing practice")
[1144] Step 5:
[1145] Information acquisition via server
[1146] Based on the received request, the server uses the video sharing platform's API to retrieve the relevant video information. Information such as the video title and URL is obtained from the API.
[1147] Input: Request
[1148] Output: Video information (e.g., title, URL)
[1149] Step 6:
[1150] Information response
[1151] The server sends the acquired video information back to the terminal.
[1152] Input: Video information
[1153] Output: Reply message (e.g., "Video title and URL")
[1154] Step 7:
[1155] Building notifications and adjusting them to your emotions
[1156] The device adjusts the content and tone of notifications based on video information received from the server and the user's emotional state.
[1157] The device uses speech synthesis technology to generate voice messages to convey to the user. For example, if the user is anxious, it might notify them in a calm tone, "Please listen calmly. The last part of the video on the basics of swing practice explained the follow-through point. The URL is http: / / example.com."
[1158] Input: Reply message, emotional state
[1159] Output: Adjusted voice message
[1160] Step 8:
[1161] Notification to the user
[1162] The device notifies the user of the generated voice message.
[1163] Input: Pre-recorded voice message
[1164] Output: Voice notification (voice message to the user)
[1165] The above outlines the specific processing steps for implementing this invention.
[1166] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1167] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1168] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1169] [Fourth Embodiment]
[1170] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1171] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1172] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1173] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1174] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1175] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1176] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1177] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1178] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1179] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1180] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1181] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1182] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1183] This invention is a system for users to quickly and accurately acquire information about specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice (overly helpful advice) at practice facilities.
[1184] User voice input
[1185] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[1186] Speech-to-text conversion and analysis
[1187] The device uses speech recognition technology to convert recorded audio into text. This converted text data is used to analyze the user's intent. Natural language processing technology is used for the analysis to identify which part of which video the user wants to know about.
[1188] Request to the server
[1189] The device sends an appropriate request to the server based on the analyzed text data. This request includes information about the video or content specified by the user. For example, a request is sent to the server to re-identify a specific YouTube video and specify the "final part" of it.
[1190] Server Processing
[1191] Based on the received request, the server retrieves relevant video information from the video sharing platform. Using existing platforms such as the YouTube API, it obtains the video's title, URL, and the content of the specified portion. The server also analyzes the video's subtitles and script to identify the information the user is seeking.
[1192] Information response
[1193] The server sends the retrieved information back to the terminal. This information includes the video title, URL, and relevant content. For example, it might provide the URL of a video titled "Basic Swing Practice" and the information that "the last part explained the key points of the follow-through."
[1194] Notification to the user
[1195] The terminal notifies the user of information received from the server via voice. This information is conveyed to the user using speech synthesis technology. As a result, the user can instantly obtain information about specific techniques and methods, and easily check the instructions needed during practice. Furthermore, it enables the user to practice independently without receiving excessive advice at the practice field.
[1196] Specific example
[1197] For example, if a user says, "Show me the last part of the swing practice video I watched yesterday by a YouTuber," the device transcribes the voice command into text and sends a request to the server based on the analysis. The server retrieves information from the YouTube video, identifies the relevant section, and sends it back to the device. The device then relays this information to the user via voice, allowing the user to improve the efficiency of their practice.
[1198] The system of this invention allows users to improve the quality of their independent practice and quickly obtain necessary information. This improves practice efficiency and makes it possible to avoid excessive intervention from others.
[1199] The following describes the processing flow.
[1200] Step 1:
[1201] The user gives voice commands using an earphone microphone. For example, they input specific instructions in voice format, such as, "Show me the last part of the swing practice video I watched yesterday again."
[1202] Step 2:
[1203] The device receives and records voice input. This recorded voice data is used for subsequent processing.
[1204] Step 3:
[1205] The device uses speech recognition technology to convert recorded audio data into text data. This is done using Google's speech recognition API, among others.
[1206] Step 4:
[1207] The device analyzes what the user wants based on the converted text data. Using natural language processing technology, it extracts information to identify a specific video and its corresponding section.
[1208] Step 5:
[1209] The device sends a request to the server based on the analysis results. The request includes the URL of the video specified by the user and information about a specific part of it.
[1210] Step 6:
[1211] The server receives the request from the terminal and uses the video sharing platform's API to retrieve the relevant video information. Specifically, it retrieves the video's title, URL, and relevant content.
[1212] Step 7:
[1213] The server further analyzes the acquired video information and uses subtitles and scripts to identify what the user is looking for. For example, it might extract specific information such as "the follow-through points explained at the end."
[1214] Step 8:
[1215] The server sends the analyzed information back to the terminal. The information sent back includes the video title, URL, and a detailed description of the relevant section.
[1216] Step 9:
[1217] The terminal constructs a message to notify the user based on information received from the server. This message is then conveyed to the user in voice format using speech synthesis technology.
[1218] Step 10:
[1219] The user receives an audio notification from their device and obtains information about a specific part of the video they were looking for. Based on this information, the user can effectively proceed with their self-practice.
[1220] (Example 1)
[1221] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1222] In conventional self-practice methods, it was difficult for users to quickly and accurately obtain information on specific techniques and methods, and there was a lack of ways to avoid excessive advice, especially at the practice field. As a result, practice efficiency decreased, and it often took a long time for users to obtain the information they needed. The present invention aims to solve these problems and provide a system that allows users to improve the quality of their self-practice.
[1223] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1224] In this invention, the server includes means for recognizing voice input from a user and converting it into text data, means for analyzing the text data to determine the user's intent, means for sending a request to the server to obtain relevant information based on the user's intent, and means for notifying the user of the information received from the server by voice. This allows the user to quickly and accurately obtain the necessary information, improve the efficiency of self-practice, and minimize the impact of excessive advice.
[1225] "Voice input" refers to voice data containing instructions or information spoken by the user.
[1226] "Text data" refers to digital information obtained by converting voice input into a string format.
[1227] "Speech recognition technology" is a technology that analyzes speech information and converts its content into text data.
[1228] "Natural language processing technology" refers to techniques that analyze text data to understand the user's intentions and emotions.
[1229] A "request" is a request sent to a server to obtain specific information.
[1230] A "server" is a computer system that receives requests via a network and provides information.
[1231] A "video sharing platform" is a web service that allows users to upload, share, and watch video content.
[1232] "Speech synthesis technology" is a technology that converts text data into speech.
[1233] "Analysis results" refer to the user's intentions and necessary information obtained through the analysis of text data.
[1234] "Video information" refers to digital information about specific video content obtained from video sharing platforms.
[1235] Modes for carrying out the invention
[1236] This invention is a system for users to quickly and accurately acquire information on specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice at the practice field. Specific embodiments are described below.
[1237] User voice input
[1238] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[1239] Speech-to-text conversion
[1240] The device converts recorded audio into text using speech recognition technologies such as the Google Cloud Speech-to-Text API. This converted text data is then used to analyze the user's intent.
[1241] Text analysis
[1242] The device analyzes the converted text data using natural language processing technologies such as the Google Cloud Natural Language API. This identifies which part of which video the user wants to know about.
[1243] Sending a request to the server
[1244] The device sends an appropriate request to the server based on the analyzed text data. This request includes information about the video or content specified by the user. For example, it might include a request to specify a particular part from a particular video sharing platform.
[1245] Server Processing
[1246] Based on the received request, the server retrieves relevant video information using the YouTube API and other tools. It obtains the video title, URL, and the content of the specified section, and analyzes the subtitles and script related to that video to identify the information the user is seeking.
[1247] Information response
[1248] The server sends the retrieved information back to the terminal. This information includes the video title, URL, and relevant content. For example, it might provide the URL of a video titled "Basic Swing Practice" and the information that "the last part explained the key points of the follow-through."
[1249] Notification to the user
[1250] The device uses speech synthesis technology, such as the Google Cloud Text-to-Speech API, to notify the user of information received from the server via voice. This allows the user to instantly obtain information about specific techniques and methods, and easily check the instructions needed during practice. It also enables the user to practice independently without receiving excessive advice at the practice field.
[1251] Specific example
[1252] For example, if a user requests to see the last part of the swing practice video from the YouTuber they watched yesterday, the following process will occur:
[1253] 1. The user gives voice commands using an earphone microphone.
[1254] 2. The device records voice commands and converts them into text using speech recognition technology.
[1255] 3. The terminal parses the text and sends a request to the server.
[1256] 4. The server uses the YouTube API to retrieve the relevant video information and analyzes the content of a specific section.
[1257] 5. The server sends the retrieved information (e.g., video URL, relevant content) back to the device.
[1258] 6. The device uses speech synthesis technology to notify the user of information.
[1259] Example of a prompt
[1260] "Show me the last part of the swing practice video I saw from the YouTuber yesterday again."
[1261] This invention allows users to quickly obtain information useful for independent practice and to continue practicing efficiently.
[1262] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1263] Step 1: Recording voice input
[1264] The user transmits voice input to the device via an earphone microphone. The input voice is recorded on the device and saved as digital audio data. A specific example of this action would be the user saying, "Show me the last part of the swing practice video I watched yesterday again."
[1265] Step 2: Text conversion using speech recognition
[1266] The device sends the recorded audio data to the Google Cloud Speech-to-Text API, which converts the audio into text data. The input is the recorded audio data, and the output is the corresponding string (e.g., "Show me the last part of the swing practice video I watched yesterday again"). Specifically, this process involves analyzing the audio waveform and generating the corresponding string.
[1267] Step 3: Analyzing text data
[1268] The device sends the acquired text data to the Google Cloud Natural Language API for analysis of the user's intent. The input is the converted text data, and the output is the analysis result, including the user's intent and request. Specifically, the NLP engine analyzes the context, keywords, and intent, and recognizes it as a request regarding the final part of a particular video.
[1269] Step 4: Generate and send the request
[1270] The terminal generates and sends a request to the server to retrieve relevant information based on the analysis results. The input is the analysis results, including the user's intent, and the output is the request data sent to the server. In a specific example, a request is generated that includes "a special video URL and section for playing the final part of the swing practice."
[1271] Step 5: Collecting video information
[1272] The server processes the received request and uses the YouTube API to collect specific video information. The input is the request data sent from the device, and the output is information including the title, URL, and content of a specific part of the relevant video. Specifically, the server accesses the YouTube API to retrieve the metadata of the video and the script or subtitles of the target section.
[1273] Step 6: Return the information
[1274] The server sends the collected video information back to the device. The input is video information obtained from the YouTube API, and the output is the information sent to the device. In a specific example, data is returned that includes the URL of a video titled "Basic Swing Practice" and information about "Follow-through points explained in the last part."
[1275] Step 7: Notify the user
[1276] The device uses text-to-speech technology (Google Cloud Text-to-Speech API) to convert information received from the server into speech and notifies the user. The input is video information sent back from the server, and the output is audio data played for the user. Specifically, the device tells the user, "In the last part of the swing practice video you watched yesterday, the key points of the follow-through were explained."
[1277] Through the steps described above, the system of the present invention can quickly and accurately provide relevant video information based on the user's voice instructions. This processing flow allows the user to effectively practice independently while avoiding excessive advice.
[1278] (Application Example 1)
[1279] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1280] The operation and maintenance procedures for machinery in factories are specialized, and employees are required to acquire this information quickly and accurately. However, finding and learning the necessary information on the job is time-consuming and laborious, and can sometimes require excessive instruction. Therefore, there is a need for systems that support efficient work and self-learning among employees.
[1281] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1282] In this invention, the server includes means for recognizing voice input and converting it into text data, means for analyzing the text data to determine the user's intent, means for sending a request to an information processing device to acquire relevant information based on the user's intent, means for notifying the user of the information received from the information processing device by voice, and means for acquiring information regarding how to operate and maintain machinery in factory work. This enables employees to efficiently acquire necessary information and engage in self-learning.
[1283] "Voice input" refers to a method in which users provide information such as commands or questions using their voice via a headset microphone or similar device.
[1284] "Text data" refers to data obtained by converting voice input into characters, and is a string of characters in a format that can be analyzed by a computer.
[1285] "Analysis" is the process of analyzing text data to understand the user's intentions and requirements.
[1286] "User intent" refers to what the user is seeking through voice input, including the content of their requests and the information they desire.
[1287] A "request" is data sent to a server to retrieve relevant information based on the user's intent.
[1288] An "information processing device" refers to a system that manages and processes various types of information and data via a network.
[1289] "Notifying" refers to the process of conveying information obtained from a server or information processing device back to the user.
[1290] "Factory work" refers to all tasks performed in workplaces such as manufacturing, using machinery and equipment.
[1291] "Operating instructions" refer to the procedures and methods for making a machine or device work correctly.
[1292] "Maintenance procedures" refer to the procedures for maintenance, inspection, and repair performed to maintain the condition of machinery and equipment.
[1293] A "video sharing platform" refers to a web service that allows users to upload, play, and share videos over the internet.
[1294] This invention's system utilizes speech recognition and natural language processing technologies to provide information on how to operate and maintain machinery in a factory. Users can input voice information using an earphone microphone and quickly and accurately obtain relevant information based on that input.
[1295] System Configuration
[1296] The system consists of the following main components:
[1297] 1. Voice Input Device: The user gives voice commands using an earphone microphone. For example, they might input a voice command such as, "Show me the last part of the maintenance video I watched before again."
[1298] 2. Speech Recognition System: This system uses speech recognition technology (such as the Google Speech Recognition API) to convert recorded audio into text data. This system operates on the device.
[1299] 3. Contextual Analysis System: Natural language processing techniques (e.g., nltk or spaCy) are used to analyze text data and determine the user's intent. At this stage, relevant information is identified based on the analysis results.
[1300] 4. Server: Receives requests based on user intent and retrieves relevant video information from video sharing platforms (such as the YouTube API). It also generates a response containing the necessary information and sends it back to the device.
[1301] 5. Notification System: The terminal notifies the user of information received from the server via voice using speech synthesis technology (e.g., Pyttsx3).
[1302] Hardware and software
[1303] The following hardware and software are used to implement this system:
[1304] Hardware: Headset with microphone, smartphone or tablet, server.
[1305] Software: Google Speech Recognition API (speech recognition), nltk or spaCy (natural language processing), YouTube API (video information retrieval), Pyttsx3 (speech synthesis).
[1306] Data processing and data calculation
[1307] 1. Speech-to-text conversion: Convert user voice input into text data using the Google Speech Recognition API.
[1308] 2. Text Analysis: The generated text data is analyzed using natural language processing libraries such as nltk and spaCy to identify the user's intent.
[1309] 3. Information Request: Based on the user's request, a request for relevant video information is made to the YouTube API, and data such as the video title and URL is retrieved.
[1310] 4. Information Notification: Use Pyttsx3 to generate the acquired information in audio format and notify the user.
[1311] Specific example
[1312] If a user gives the voice command, "Show me the last part of the robot maintenance video I saw before again," the system will perform the following actions:
[1313] 1. Convert voice input to text.
[1314] 2. Analyze the converted text data to identify what information the user is looking for.
[1315] 3. Use the YouTube API to retrieve relevant video information.
[1316] 4. The acquired information (video title, URL, specified portion) is generated using speech synthesis technology and notified to the user.
[1317] Example of a prompt
[1318] User: "Show me the last part of the robot maintenance video I saw before again."
[1319] This allows users to efficiently obtain necessary information and streamline factory operations. The system reduces excessive on-site instruction and supports employees' self-directed learning.
[1320] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1321] Step 1:
[1322] The user gives voice commands using an earphone microphone. For example, they might input a voice command such as, "Show me the last part of the robot maintenance video I watched before again." The input is the user's voice data.
[1323] Step 2:
[1324] The device uses speech recognition technology (Google Speech Recognition API) to convert the user's voice data into text data. In this process, voice data is provided as input, and the converted text data is output.
[1325] Step 3:
[1326] The device uses natural language processing technology (nltk or spaCy) to analyze the generated text data and identify the user's intent. In this step, text data is the input, and the analysis outputs the user's intent (e.g., a specific part of a video to play).
[1327] Step 4:
[1328] The terminal sends a request to the server to retrieve relevant information based on the user's intent. In this step, data including the user's intent is input, and the request data to the server is output.
[1329] Step 5:
[1330] Based on the received request, the server retrieves relevant video information from the video sharing platform (YouTube API). In this step, the request data is used as input, and the video title, URL, and specific information are output.
[1331] Step 6:
[1332] The server sends the acquired information back to the terminal. In this step, video information is used as input and is sent to the terminal.
[1333] Step 7:
[1334] The terminal uses speech synthesis technology (Pyttsx3) to notify the user of the information received from the server via voice. In this step, the received video information is the input, and the output is an audio notification.
[1335] Step 8:
[1336] The user receives an audio notification from their device and reconfirms information on how to operate or maintain the relevant machine. In this step, the audio notification is provided as input, and an action is output that plays the information the user is requesting.
[1337] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1338] This invention is a system for users to quickly and accurately acquire information about specific techniques and methods when practicing independently. This system recognizes the user's voice input and provides appropriate video information, thereby improving the efficiency of independent learning and minimizing the impact of excessive advice (overly helpful advice) at practice facilities. Furthermore, this invention can further enhance the user experience by incorporating an emotion engine that recognizes the user's emotions.
[1339] User voice input
[1340] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[1341] Speech-to-text conversion and emotion recognition
[1342] The device uses speech recognition technology to convert recorded speech into text. This utilizes a speech recognition API. The converted text data is used to analyze the user's intent. Simultaneously, an emotion engine is used to analyze the user's emotions from the speech. The emotion engine analyzes the tone and characteristics of the speech to determine the user's emotional state.
[1343] User intent analysis and request generation
[1344] The device analyzes the user's intent based on text data and sentiment data. Combining natural language processing and sentiment analysis technologies, it extracts information to identify a specific video and its associated sections. Based on the analysis results, it sends an appropriate request to the server. This request includes information about the video and content specified by the user.
[1345] Server Processing
[1346] Based on the request received from the terminal, the server uses the video sharing platform's API to retrieve relevant video information. Specifically, it obtains the video's title, URL, and relevant content. Furthermore, the server analyzes the subtitles and script related to the video to identify the information the user is seeking.
[1347] Information response and notification
[1348] The server sends the analyzed information back to the device. This information includes the video title, URL, and a detailed description of the relevant parts. Based on the information received from the server, the device constructs a message to notify the user. Using speech synthesis technology, this message is delivered to the user in voice format. The emotion engine adjusts the content and tone of the notification according to the user's emotional state. For example, if the user is anxious, the information is delivered in a calm tone; conversely, if the user is relaxed, the information is delivered in a more casual tone.
[1349] Specific example
[1350] For example, if a user says, "Show me the last part of the YouTuber's swing practice video I watched yesterday," the device transcribes the voice command into text and then uses an emotion engine to analyze the user's emotional state. Based on the analyzed text data and emotion data, the device identifies the user's intent and sends a request to the server. The server retrieves information from the YouTube video, identifies the relevant section, and sends it back to the device. The device then uses this information, along with the emotion engine's analysis results, to send a voice notification. For example, if the user is anxious, the notification might say, "Please listen calmly. The last part of the video about the basics of swing practice explained the key points of the follow-through. The URL is http: / / example.com."
[1351] The system of this invention allows users to improve the quality of their independent practice and quickly obtain necessary information. This improves practice efficiency and allows them to avoid excessive intervention from others. Furthermore, by considering the user's emotional state, it can provide a more human-like interaction.
[1352] The following describes the processing flow.
[1353] Step 1:
[1354] The user gives voice commands using an earphone microphone. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again."
[1355] Step 2:
[1356] The device records voice input. This recorded voice data is used for speech recognition processing.
[1357] Step 3:
[1358] The device uses speech recognition technology to convert recorded audio data into text data. It utilizes a speech recognition API to accurately transcribe the user's speech into text.
[1359] Step 4:
[1360] The device analyzes the converted text data to identify the user's intent. Using natural language processing technology, it determines the video or specific part of the video the user is looking for.
[1361] Step 5:
[1362] The device simultaneously uses an emotion engine to analyze the user's emotional state from the voice data. It analyzes the tone of voice and speaking style to determine whether the user is excited, calm, or otherwise emotional.
[1363] Step 6:
[1364] The device sends an appropriate request to the server based on the analyzed text and sentiment data. This request includes information about the video identified by the user and the content of relevant parts of it.
[1365] Step 7:
[1366] Based on the request received from the terminal, the server uses the video sharing platform's API to retrieve the relevant video information. Specifically, it retrieves the video's title, URL, and relevant content.
[1367] Step 8:
[1368] The server further analyzes the acquired video information to identify detailed explanations of relevant parts. For example, it extracts the content explained in the final part of the video from the script or subtitles.
[1369] Step 9:
[1370] The server returns the analyzed information to the terminal. This information includes the video title, URL, and a detailed description of the relevant section.
[1371] Step 10:
[1372] The terminal constructs a message to notify the user based on information received from the server. The message is created using speech synthesis technology and adjusted according to the user's emotional state, which is analyzed by the emotion engine.
[1373] Step 11:
[1374] The device provides voice notifications and conveys acquired information to the user. For example, it might notify the user of the content and URL explained at the end of a video titled "Basic Swing Practice" in a calm or casual tone, depending on the user's emotional state.
[1375] Step 12:
[1376] Users receive audio notifications from their devices and obtain information about specific parts of the video they need. This allows users to effectively practice on their own.
[1377] (Example 2)
[1378] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1379] Conventional self-training and learning support systems have struggled to provide appropriate information quickly and accurately based on user voice input. Furthermore, they fail to consider the user's emotional state, resulting in a poor user experience. To solve these problems, the present invention aims to provide a system that efficiently transcribes user voice instructions into text, performs emotion analysis, and provides appropriate video information.
[1380] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1381] In this invention, the server includes means for recognizing voice input from a user and converting it into text data; means for analyzing the user's emotions from the text data and voice data; means for determining the user's intentions by analyzing the text data and emotion data; means for sending a request to the server to obtain relevant information based on the user's intentions; means for analyzing information obtained from the server through a video sharing platform and identifying the parts requested by the user; means for notifying the user of the information received from the server by voice; and means for adjusting the content and tone of the notification according to the user's emotional state. This enables the user to quickly and accurately obtain the necessary information and receive feedback that is appropriate to their emotions.
[1382] A "user" is an individual or organization that is using the system to obtain information.
[1383] "Voice input" refers to voice data used by a user to give instructions to a system via an earphone microphone or other voice input device.
[1384] "Text data" refers to character information converted from voice input using speech recognition technology.
[1385] "Emotional data" refers to information obtained by analyzing the user's emotional state from voice data.
[1386] "Speech recognition technology" is a technology that converts voice input into text data, and typically uses APIs or specialized software.
[1387] "Emotional analysis technology" is a technology that analyzes a user's emotional state from voice data, analyzing the tone of voice and characteristics of their speaking style.
[1388] "Natural language processing technology" is a technology that analyzes text data to understand the user's intent.
[1389] A "request" is a request for information to be retrieved, sent to a server based on the user's intent.
[1390] A "server" is a computer system that processes information based on requests and provides it to users.
[1391] A "video sharing platform" is a website or service that hosts videos that are accessible to users.
[1392] "Subtitles" are information used to display the content of a video in text format.
[1393] A "script" is text data that records the content and dialogue of a video.
[1394] "Speech synthesis technology" is a technology that converts text data into a speech format for playback.
[1395] "Notification content" refers to the specific information conveyed to the user based on information obtained from the server.
[1396] "Tone" refers to the emotion or attitude expressed in the audio of a notification.
[1397] "User intent" refers to the purpose or desired information that the user is trying to convey to the system through voice commands.
[1398] This invention is a system designed to enable users to quickly and accurately acquire information about specific techniques and methods when practicing independently. It improves the efficiency of independent learning by recognizing user voice input and providing appropriate video information. Furthermore, it can enhance the user experience by recognizing user emotions. This system includes three main elements: the user, the terminal, and the server.
[1399] Hardware and software to be used
[1400] 1. Earphone microphone (hardware): Used by the user to give voice commands.
[1401] 2. Device (hardware and software): A device for recording voice input, converting it to text, performing sentiment analysis, and intent analysis. The Google Cloud Speech-to-Text API is used for text data analysis, and the Azure Cognitive Services Emotion API is used for sentiment analysis.
[1402] 3. Server (hardware and software): Based on user instructions, it retrieves video information, analyzes that information, and identifies the necessary parts. The YouTube Data API is used for this purpose.
[1403] System operation
[1404] The user gives voice commands using an earphone microphone. For example, they might give specific instructions such as, "Show me the last part of the swing practice video I watched yesterday again." This voice is recorded by the device.
[1405] The device converts recorded audio into text data using the Google Cloud Speech-to-Text API. Simultaneously, it analyzes the user's emotional state from the audio data using the Azure Cognitive Services Emotion API. This text data and emotional data are then used for further analysis.
[1406] The device analyzes the user's intent using OpenAI's GPT-4 model based on text and sentiment data. Based on the analysis results, the specific video and its corresponding section that the user requests are identified. A request containing this information is then sent to the server.
[1407] The server uses the YouTube Data API to search for the specified video information and retrieves the video's title, URL, content, etc. The server then analyzes the video's subtitles and script to identify the parts requested by the user.
[1408] The server sends the analyzed information back to the device. The device uses the Google Cloud Text-to-Speech API to generate a voice message based on the information received from the server. At the same time, it adjusts the tone of the notification based on the sentiment analysis results. For example, if the user is anxious, the information will be delivered in a calm tone.
[1409] The device plays the generated voice message to the user. This allows the user to quickly and accurately obtain the necessary information.
[1410] Specific example
[1411] For example, if a user instructs, "Show me the last part of the swing practice video I watched yesterday again," the device records this audio and transcribes it using the Google Cloud Speech-to-Text API. Next, it performs sentiment analysis using the Azure Cognitive Services Emotion API. The analyzed text and sentiment data are further analyzed using the OpenAI GPT-4 model to identify a specific video segment. Based on this information, a request is sent to the server, and the relevant video information is retrieved using the YouTube Data API. The retrieved information is sent back to the device, and an audio message is generated using the Google Cloud Text-to-Speech API. The device then plays this message to the user.
[1412] For example, the user might prompt with the following message: "Show me the last part of the swing practice video I watched yesterday by a YouTuber." Based on this, the system operates and provides the user with the necessary video information.
[1413] This allows users to improve the quality of their self-practice and obtain necessary information quickly and accurately. Furthermore, feedback that takes into account the user's emotional state can provide a more human-like interaction.
[1414] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1415] Step 1:
[1416] The user gives voice commands using an earphone microphone. These voice commands are specific, such as, "Show me the last part of the swing practice video I watched yesterday again." The device records this voice and prepares for the next processing step.
[1417] Input: User's voice instructions
[1418] Output: Recorded audio data
[1419] Step 2:
[1420] The device converts the recorded audio data into text data using the Google Cloud Speech-to-Text API. The text data is the extracted content of the voice instructions as written information. The device also analyzes the user's emotional state from the audio data using the Azure Cognitive Services Emotion API.
[1421] Input: Recorded audio data
[1422] Output: Text data and sentiment data
[1423] Step 3:
[1424] The device analyzes the user's intent using OpenAI's GPT-4 model based on acquired text and sentiment data. Specifically, it identifies the particular video section the user is looking for.
[1425] Input: Text data and sentiment data
[1426] Output: Analyzed intent data (information from specific parts of the video)
[1427] Step 4:
[1428] The device sends a request to the server to retrieve appropriate information based on the analyzed intent data. The request to the server includes information about the video and content specified by the user. The request may include information about the title or specific portion of a particular video.
[1429] Input: Intent data
[1430] Output: Request to send to the server
[1431] Step 5:
[1432] Based on the request received from the terminal, the server uses the API of a video sharing platform (e.g., YouTube) to search for and retrieve the relevant video information. Specifically, it obtains the video title, URL, and relevant content, and further analyzes subtitles and scripts to identify the requested portion.
[1433] Input: Request from terminal
[1434] Output: Acquired video information (title, URL, subtitles, content of specific section)
[1435] Step 6:
[1436] The server sends the acquired video information back to the device. Based on the received information, the device generates an audio notification message using the Google Cloud Text-to-Speech API. This message includes information about the specific part of the video requested by the user. It also adjusts the tone of the notification based on the user's sentiment data.
[1437] Input: Acquired video information and emotion data
[1438] Output: Generated voice notification message
[1439] Step 7:
[1440] The device plays a generated voice notification message to the user. For example, if the user is anxious, it will notify the user in a calm tone, saying, "Please listen calmly. The last part of the video on the basics of swing practice explained the key points of the follow-through. The URL is http: / / example.com." Through voice notifications, the user can quickly and accurately obtain the information they need.
[1441] Input: Voice notification message
[1442] Output: Voice notification to the user
[1443] This allows users to improve the quality of their self-practice and quickly and accurately obtain the information they need. Furthermore, by considering the user's emotional state, a better user experience can be provided.
[1444] (Application Example 2)
[1445] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1446] While conventional learning support systems could retrieve relevant information based on user voice commands, they struggled to provide information appropriately according to the user's emotional state. This resulted in insufficient improvement in user learning efficiency and experience. Furthermore, particularly in content distribution services, there was a lack of means to quickly retrieve the learning content learners desired and to provide appropriate interactions that responded to the user's emotions.
[1447] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1448] In this invention, the server includes means for sending a request to the server to obtain relevant information based on the user's intent; means for notifying the user of the information received from the server by voice; means for determining the user's emotional state using emotion analysis technology; means for adjusting the content and tone of the notification based on the information received from the server and the user's emotional state; and means for learners to obtain learning content through voice input in a content distribution service. This enables the provision of information appropriately according to the user's emotional state, improving learning efficiency and enabling more human-like interaction.
[1449] "Voice input" refers to audio data emitted by a user using a voice collection device such as a microphone.
[1450] "Text data" refers to data obtained by converting voice input into text information using technologies such as speech recognition.
[1451] A "request" refers to the content of a request sent to a server to retrieve relevant information based on the user's intent or needs.
[1452] A "server" refers to a computer or computer system that stores and processes information over a network and provides that information to clients.
[1453] "Speech recognition technology" is a technology that analyzes speech input collected by microphones or other devices and converts it into text data.
[1454] A "video sharing platform" refers to an online service that allows users to upload, share, and watch video content over the internet.
[1455] "Emotion analysis technology" is a technology that analyzes data such as voice and text to estimate a user's emotional state.
[1456] "Notification" refers to the act of a system transmitting information or messages to a user.
[1457] "Tone" refers to the tone of voice and speaking style characteristics in messages such as notifications.
[1458] A "content distribution service" refers to a service that delivers digital content such as videos and music to users via the internet.
[1459] "Learning content" refers to educational videos, texts, and other multimedia materials that users use for learning.
[1460] The system for implementing this invention uses the following hardware and software.
[1461] Hardware and software to be used
[1462] Microphone (sound acquisition device):
[1463] It is used to collect voice input from users.
[1464] Device (smartphone, etc.):
[1465] This device performs speech recognition and emotion analysis technologies, processes data, and notifies the user of the results.
[1466] Speech recognition API:
[1467] To convert voice input into text data, you can use, for example, Google's speech recognition API.
[1468] Generative AI models:
[1469] To perform emotion analysis, for example, we can use the pipeline from the "transformers" library.
[1470] server:
[1471] The system uses the API of a video sharing platform (e.g., the YouTube Data API) to retrieve relevant video information and sends that information back to the device.
[1472] Speech synthesis technology:
[1473] This is used to notify users of information from the server as voice messages.
[1474] Explanation of the process
[1475] 1. Collection and recognition of voice input:
[1476] The user provides voice input via a microphone. For example, they might give a command like, "Show me the last part of the swing practice video I watched yesterday again."
[1477] The device collects this audio and uses speech recognition technology to convert it into text data.
[1478] 2. Emotional analysis and intention judgment:
[1479] The device uses a generative AI model to analyze the user's emotional state from text data. In this process, it uses a generative AI model (for example, the "transformers" library's pipeline) to identify emotions.
[1480] The device uses speech recognition technology to analyze text data and determine the user's intent. For example, it can understand that the user is requesting to play a specific video.
[1481] 3. Sending requests and retrieving information:
[1482] The device sends a request to the server to retrieve relevant information based on the user's intent. This request includes the search query for the required video.
[1483] The server uses an API to retrieve relevant video information (title, URL, etc.) from the video sharing platform.
[1484] 4. Information response and notification:
[1485] The server sends the acquired video information back to the terminal.
[1486] The device adjusts the content and tone of notifications based on the information it receives and the user's emotional state.
[1487] The device uses speech synthesis technology to generate an appropriate voice message and notify the user.
[1488] Specific example
[1489] For example, if a user instructs, "Show me the part of the math lesson from yesterday where we saw the proof of Fermat's Little Theorem again":
[1490] 1. User voice collection and recognition:
[1491] The user's voice commands are collected by the device and converted into text data using speech recognition technology, such as "Show me again the part of the math lesson yesterday that showed the proof of Fermat's Little Theorem."
[1492] 2. Emotional analysis and intention judgment:
[1493] Text data is analyzed using a generation AI model to identify the user's emotional state.
[1494] The device determines that the user is requesting the playback of a specific math video.
[1495] 3. Sending requests and retrieving information:
[1496] Based on user instructions, a request is sent to the server, and relevant mathematical video information is retrieved from the video sharing platform.
[1497] 4. Information response and notification:
[1498] Based on information received from the server, the device adjusts the tone of notifications according to the user's emotional state.
[1499] For example, an audio message is generated and sent to the user saying, "Here is the video proof of Fermat's Little Theorem that you requested. The URL is http: / / example.com."
[1500] Example of a prompt
[1501] Perform sentiment analysis on the AI model using the following prompt:
[1502] "Please analyze the emotion in the following sentence: 'Show me again the part of the math lesson yesterday that showed the proof of Fermat's Little Theorem.'"
[1503] This embodiment allows users to quickly obtain the necessary learning materials and receive appropriate support tailored to their emotional state.
[1504] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1505] Step 1:
[1506] Collection of user voice input
[1507] The user uses a microphone to provide voice input. For example, they might say, "Show me the last part of the swing practice video I watched yesterday again."
[1508] Input: User's voice data
[1509] Output: Audio data is collected.
[1510] Step 2:
[1511] Text conversion of audio data
[1512] The device uses a speech recognition API to convert audio data into text data. Specifically, it uses speech recognition technology (for example, Google's speech recognition API) to convert speech into text information.
[1513] Input: Collected audio data
[1514] Output: Text data (Example: "Show me the last part of the swing practice video I watched yesterday again.")
[1515] Step 3:
[1516] Emotion analysis
[1517] The device uses a generative AI model to analyze the user's emotional state from text data. It uses emotion analysis techniques (e.g., the transformers library's pipeline) to analyze the tone and characteristics of speech to determine the emotional state.
[1518] Input: Text data
[1519] Output: Emotion analysis results (e.g., "anxiety")
[1520] Step 4:
[1521] Intent analysis and request generation
[1522] The device analyzes the user's intent based on text data and sentiment analysis results. Specifically, it uses natural language processing technology to identify what the user is looking for.
[1523] Based on the analysis results, the terminal sends a request to the server to retrieve relevant information.
[1524] Input: Text data, sentiment analysis results
[1525] Output: Request (Example: "Retrieve video information for swing practice")
[1526] Step 5:
[1527] Information acquisition via server
[1528] Based on the received request, the server uses the video sharing platform's API to retrieve the relevant video information. Information such as the video title and URL is obtained from the API.
[1529] Input: Request
[1530] Output: Video information (e.g., title, URL)
[1531] Step 6:
[1532] Information response
[1533] The server sends the acquired video information back to the terminal.
[1534] Input: Video information
[1535] Output: Reply message (e.g., "Video title and URL")
[1536] Step 7:
[1537] Building notifications and adjusting them to your emotions
[1538] The device adjusts the content and tone of notifications based on video information received from the server and the user's emotional state.
[1539] The device uses speech synthesis technology to generate voice messages to convey to the user. For example, if the user is anxious, it might notify them in a calm tone, "Please listen calmly. The last part of the video on the basics of swing practice explained the follow-through point. The URL is http: / / example.com."
[1540] Input: Reply message, emotional state
[1541] Output: Adjusted voice message
[1542] Step 8:
[1543] Notification to the user
[1544] The device notifies the user of the generated voice message.
[1545] Input: Pre-recorded voice message
[1546] Output: Voice notification (voice message to the user)
[1547] The above outlines the specific processing steps for implementing this invention.
[1548] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1549] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1550] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1551] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1552] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1553] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1554] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1555] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1556] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1557] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1558] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1559] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1560] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1561] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1562] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1563] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1564] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1565] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1566] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1567] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1568] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1569] The following is further disclosed regarding the embodiments described above.
[1570] (Claim 1)
[1571] A means of recognizing voice input from a user and converting it into text data,
[1572] A means for analyzing the aforementioned text data to determine the user's intent,
[1573] A means for sending a request to the server to obtain relevant information based on the user's intent,
[1574] A means for notifying the user of the information received from the aforementioned server via voice,
[1575] A system that includes this.
[1576] (Claim 2)
[1577] The system according to claim 1, wherein the means for recognizing the voice input and converting it into text data utilizes voice recognition technology.
[1578] (Claim 3)
[1579] The system according to claim 1, wherein the server includes means for acquiring video information from a video sharing platform.
[1580] "Example 1"
[1581] (Claim 1)
[1582] A means of recognizing voice input from a user and converting it into text data,
[1583] A means for analyzing the aforementioned text data to determine the user's intent,
[1584] A means for sending a request to the server to obtain relevant information based on the user's intent,
[1585] A means for notifying the user of the information received from the aforementioned server via voice,
[1586] A system that includes this.
[1587] (Claim 2)
[1588] The system according to claim 1, wherein the means for recognizing the voice input and converting it into text data utilizes voice recognition technology.
[1589] (Claim 3)
[1590] The system according to claim 1, wherein the server includes means for acquiring video information from a video sharing platform.
[1591] (Claim 4)
[1592] The system according to claim 1, comprising means for using natural language processing techniques to analyze the aforementioned text data.
[1593] (Claim 5)
[1594] The system according to claim 1, further comprising means for sending a request to a server based on the results of analyzing the aforementioned text data.
[1595] (Claim 6)
[1596] The system according to claim 1, wherein the server includes means for analyzing the content of a specific portion of the acquired video information.
[1597] (Claim 7)
[1598] The system according to claim 1, further comprising means for notifying a user of information from the server using speech synthesis technology.
[1599] "Application Example 1"
[1600] (Claim 1)
[1601] A means of recognizing voice input and converting it into text data,
[1602] A means for analyzing the aforementioned text data to determine the user's intent,
[1603] Means for sending a request to an information processing device to obtain relevant information based on the user's intent,
[1604] A means for notifying the user of the information received from the information processing device via voice,
[1605] A means of obtaining information on how to operate and maintain machinery in factory work,
[1606] A system that includes this.
[1607] (Claim 2)
[1608] The system according to claim 1, wherein the means for recognizing the voice input and converting it into text data utilizes voice recognition technology.
[1609] (Claim 3)
[1610] The system according to claim 1, wherein the information processing device includes means for acquiring video information from a video sharing platform.
[1611] "Example 2 of combining an emotion engine"
[1612] (Claim 1)
[1613] A means of recognizing voice input from a user and converting it into text data,
[1614] A means for analyzing the user's emotions from the aforementioned text data and audio data,
[1615] A means for analyzing the aforementioned text data and sentiment data to determine the user's intent,
[1616] A means for sending a request to the server to obtain relevant information based on the user's intent,
[1617] A means for analyzing information obtained from the aforementioned server through a video sharing platform and identifying the portion requested by the user,
[1618] A means for notifying the user of the information received from the aforementioned server via voice,
[1619] A means of adjusting notification content and tone according to the user's emotional state,
[1620] A system that includes this.
[1621] (Claim 2)
[1622] The system according to claim 1, wherein the means for recognizing the voice input and converting it into text data utilizes voice recognition technology.
[1623] (Claim 3)
[1624] The system according to claim 1, wherein the server includes means for acquiring video information from a video sharing platform.
[1625] (Claim 4)
[1626] The system according to claim 1, wherein the server includes means for analyzing subtitles or scripts related to the video.
[1627] "Application example 2 when combining with an emotional engine"
[1628] (Claim 1)
[1629] A means of recognizing voice input from a user and converting it into text data,
[1630] A means for analyzing the aforementioned text data to determine the user's intent,
[1631] A means for sending a request to the server to obtain relevant information based on the user's intent,
[1632] A means for notifying the user of the information received from the aforementioned server via voice,
[1633] A means of determining a user's emotional state using emotion analysis technology,
[1634] A means for adjusting the content and tone of notifications based on information received from the server and the user's emotional state,
[1635] In a content distribution service, a means by which learners acquire learning content through voice input,
[1636] A system that includes this.
[1637] (Claim 2)
[1638] The system according to claim 1, wherein the means for recognizing the voice input and converting it into text data utilizes voice recognition technology.
[1639] (Claim 3)
[1640] The system according to claim 1, wherein the server includes means for acquiring video information from a video sharing platform. [Explanation of Symbols]
[1641] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of recognizing voice input from a user and converting it into text data, A means for analyzing the aforementioned text data to determine the user's intent, A means for sending a request to the server to obtain relevant information based on the user's intent, A means for notifying the user of the information received from the aforementioned server via voice, A system that includes this.
2. The system according to claim 1, wherein the means for recognizing the voice input and converting it into text data utilizes voice recognition technology.
3. The system according to claim 1, wherein the server includes means for acquiring video information from a video sharing platform.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A