System
The smart earphone system addresses the limitations of AR glasses by using lightweight, battery-efficient devices to capture and analyze audio and video data, enabling real-time information access and emergency responses.
Patent Information
- Application Number
- JP2024119143
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
Smart Images

Figure 2026018082000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional AR glasses and other wearable devices have numerous drawbacks, such as being large and heavy, being sensitive to light, and having short battery life, making them unsuitable for everyday use. Furthermore, they lack the ability to quickly access specific information when needed, making it difficult to improve the user experience. The present invention aims to solve these problems and enable users to obtain the information they need in real time, thereby making daily life more comfortable and efficient. [Means for solving the problem]
[0005] To solve the above problems, the present invention provides the following means: a system including means for capturing a user's voice, means for capturing a video of the user's viewpoint, means for transmitting the captured voice data and video data to a server, means for the server to analyze the voice data and recognize the user's intention, means for the server to analyze the video data and extract necessary information, means for generating a response based on the analyzed data and transmitting the result data to a terminal, and means for notifying the user of the transmitted result. This system allows users to obtain necessary information in real time using lightweight, battery-efficient smart earphones, while eliminating the drawbacks of AR glasses.
[0006] "User" refers to a person who uses this system to obtain information or receive services.
[0007] "Audio capture means" refers to a device or software that captures the audio produced by a user and processes it as digital data.
[0008] "Means for capturing viewpoint images" refers to a device or software for recording images in the direction a user is looking as digital data.
[0009] "Server" refers to a computer system that analyzes audio and video data, extracts necessary information, and processes it.
[0010] "Analysis" refers to the process of analyzing received data using computational processes and algorithms to convert it into meaningful information.
[0011] "Means for recognizing intent" refers to technology that understands the user's intent and requests from the content of their speech and generates an appropriate response.
[0012] "Means for extracting necessary information" refers to techniques for identifying and extracting information relevant to the purpose from the captured data.
[0013] "Means for generating a response" refers to the process of creating data or a message to respond to the user based on the analysis results.
[0014] A "terminal" refers to a device such as a smart earphone worn by a user, which communicates with the server and provides information to the user.
[0015] "Means for notifying" refers to a device or software for notifying the user of the results received from the server by voice or other means.
[0016] "Results Data" refers to the final information that is analyzed and processed by the Server and provided to the User. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system has a series of functions: capturing audio, capturing viewpoint video, analyzing data, generating results, and notifying the user of those results. Specific embodiments of the system are described below.
[0039] System configuration
[0040] This system consists of the following main components:
[0041] 1. User device (smart earphones and first-person camera)
[0042] 2. Server (AI processing unit)
[0043] 3. Internet connection
[0044] User terminal
[0045] The user terminal includes a microphone for capturing the user's voice and a first-person camera for capturing the user's point of view. The terminal also includes a communication module for transmitting data to a server.
[0046] server
[0047] The server includes an AI processing unit for analyzing the received audio and video data. This allows the server to recognize the user's intent and execute processing to provide the necessary information. The server also returns the resulting data to the device, which notifies the user of the results.
[0048] System Operation
[0049] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[0050] Audio capture and transmission
[0051] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and generates audio data. At the same time, a camera capturing the user's point of view captures video of the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed, and then sent to a server via the Internet.
[0052] Server parsing and processing
[0053] The server analyzes the received voice data and recognizes the user's intent. For example, it extracts the intent "signboard" and "please translate" from the voice. At the same time, it analyzes the video data and identifies the text area on the signboard. The identified text information is translated into the specified language by the translation engine.
[0054] Generate and return result data
[0055] The server generates result data including the translation result, compresses this data again, and transmits it to the terminal.
[0056] User Notification
[0057] The device decompresses the received result data and notifies the user by voice, for example, "This sign says 'Business hours: 10:00-22:00'."
[0058] Specific examples
[0059] Translation request example
[0060] 1. User: "I want the text on the sign translated."
[0061] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[0062] 3. Server: The speech recognition engine converts the voice data into text, the video analysis engine recognizes the characters on the sign, and the translation engine then translates the characters into the specified language.
[0063] 4. Server: Sends the translation results to the device.
[0064] 5. Terminal: The translation result is notified to the user by voice.
[0065] Anomaly detection example
[0066] 1. User: (Unconscious, collapsed)
[0067] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[0068] 3. Server: Analyzes abnormal situations and determines that the user is in danger.
[0069] 4. Server: Contacts emergency services and provides them with the necessary information, and also notifies emergency contacts.
[0070] 5. Terminal: The user is notified by voice that an ambulance has been dispatched.
[0071] In this way, the system of the present invention can provide users with information quickly and accurately in their daily lives and can also respond to emergencies.
[0072] The processing flow will be explained below.
[0073] Specific processing steps of the program
[0074] Example 1: Processing a translation request
[0075] Step 1:
[0076] The user speaks to the smart earphones and says, "I want the text on the sign translated."
[0077] Step 2:
[0078] The device's microphone captures the user's voice, which is then converted into digital data.
[0079] Step 3:
[0080] The device's camera captures an image of the sign from the user's point of view, and the image data is also converted into digital data.
[0081] Step 4:
[0082] The audio and video data captured by the device is temporarily stored in local storage.
[0083] Step 5:
[0084] The terminal transmits audio and video data to the server using a secure protocol.
[0085] Step 6:
[0086] The server receives the voice data and uses a speech recognition engine to convert the digital speech into text, for example, recognizing a command such as "Please translate the text on the sign."
[0087] Step 7:
[0088] The server receives the video data and uses computer vision techniques to identify text areas in the video, for example, to extract text on a sign.
[0089] Step 8:
[0090] The server combines the results of voice recognition and video analysis to understand the user's intent (in this case, "translating the text on the sign").
[0091] Step 9:
[0092] The server uses a translation engine to translate the sign text into the specified language, for example, translating Japanese text into English.
[0093] Step 10:
[0094] The server generates result data including the translation result, compresses it, and transmits it to the terminal.
[0095] Step 11:
[0096] The terminal receives and decompresses the resulting data.
[0097] Step 12:
[0098] The device will then notify the user of the translation results by voice, for example, saying, "This sign says 'Business hours: 10:00-22:00'."
[0099] Example 2: Emergency call due to abnormality detection
[0100] Step 1:
[0101] The user falls into an abnormal state such as losing consciousness.
[0102] Step 2:
[0103] The device's camera captures the user's unusual postures and situations.
[0104] Step 3:
[0105] The device acquires anomaly detection data from biometric sensors and other devices and temporarily stores it in local storage.
[0106] Step 4:
[0107] The video data and anomaly detection data captured by the device are sent to the server using a secure protocol.
[0108] Step 5:
[0109] The server analyzes the received video data and anomaly detection data.
[0110] Step 6:
[0111] The server determines from the analysis results that the user is in danger.
[0112] Step 7:
[0113] The server automatically contacts the nearest emergency services and provides them with the necessary information, such as the user's current location and condition.
[0114] Step 8:
[0115] The server generates data to notify the registered emergency contacts and transmits this to the terminal.
[0116] Step 9:
[0117] The device receives the emergency call results and provides the information to relatives and friends via voice or other notification means.
[0118] Step 10:
[0119] The terminal notifies the user by voice that the ambulance has been dispatched, for example, by saying, "The ambulance has been dispatched. Please wait a moment."
[0120] The above are the specific processing steps in the system of the present invention.
[0121] Example 1
[0122] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0123] Current speech recognition and translation systems can sometimes make it difficult for users to quickly and accurately obtain the information they need. Furthermore, there is a lack of systems that can monitor users' status in real time and provide appropriate responses in emergencies. Therefore, there is a need for systems that can provide users with quick and accurate information in everyday life and respond appropriately in emergencies.
[0124] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0125] In this invention, the server includes means for analyzing voice data using a voice recognition engine to recognize the user's intention, means for analyzing video data using image analysis means to extract necessary information, and means for translating the extracted information into a specified language using a translation engine, thereby making it possible to recognize the user's intention from the voice data, extract necessary information from the video data, and further translate it.
[0126] An "audio capturing means" is a device or software for collecting and digitizing a user's voice.
[0127] The "means for capturing viewpoint images" is a device or software for collecting images based on the user's viewpoint.
[0128] A "compression means" is a device or software that compresses audio and video data to reduce its size.
[0129] The "means for transmitting to the server" is a device or software for transmitting compressed audio data and video data to the server via the Internet.
[0130] "Means for analyzing and recognizing the user's intent using a voice recognition engine" refers to software that analyzes voice data on a server, converts it into text data, and understands the user's requests and intent.
[0131] "Means for analyzing and extracting necessary information using image analysis means" refers to software that analyzes video data on a server and identifies and extracts important information from it.
[0132] The "means for translating into a specified language using a translation engine" is software for translating extracted information into another language.
[0133] The "means for generating a response and transmitting the resulting data to the terminal" refers to software and communication devices for generating a response based on the analyzed and translated data and transmitting the resulting data to the terminal.
[0134] The "means for notifying the user using a voice synthesis engine" refers to software and a playback device for converting the result data into voice data and notifying the user.
[0135] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system has a series of functions: capturing audio, capturing viewpoint video, analyzing data, generating results, and notifying the user of those results. Specific embodiments of the system are described below.
[0136] System configuration
[0137] This system consists of the following main components:
[0138] 1. User device (smart earphones and first-person camera)
[0139] 2. Server (AI processing unit)
[0140] 3. Internet connection
[0141] User terminal
[0142] The user terminal includes a microphone to capture the user's voice and a first-person camera to capture the user's point of view. The terminal also has a communications module for transmitting data to a server. A digital microphone is used to capture the voice, and a high-resolution camera is used to capture the point of view. Data compression is performed using MP3 format for audio data and JPEG or MPEG format for video data.
[0143] server
[0144] The server includes an AI processing unit for analyzing the received audio and video data. For example, Google Speech-to-Text is used as the speech recognition engine, and OpenCV is used for image analysis. Furthermore, the Google Translate API is used as the translation engine. The analyzed and translated data is stored in JSON format and sent back to the device after GZIP compression. The device then uses a speech synthesis engine (e.g., Google Text-to-Speech) to notify the user.
[0145] System Operation
[0146] In an embodiment of the system, when a user speaks to the smart earphones, the system operates as follows: When the user speaks to the smart earphones, saying, "I want the words on the sign translated," the digital microphone on the user device captures this voice and generates audio data. At the same time, the first-person camera captures video of the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed into MP3 and JPEG formats, and then transmitted to a server via the Internet.
[0147] The server analyzes the received voice data using Google Speech-to-Text to recognize the user's intent. For example, it extracts the intents "signboard" and "translate" from the voice. At the same time, it analyzes the video data using OpenCV to identify the text area on the signboard. The identified text information is translated into the specified language using the Google Translate API.
[0148] The server generates result data containing the translation results, composes this data in JSON format, compresses it with GZIP, and sends it to the device. The device then decompresses the received result data and uses Google Text-to-Speech to notify the user aloud. For example, "This sign says 'Business hours: 10:00-22:00'."
[0149] Specific examples
[0150] Translation request example
[0151] 1. User: "I want the text on the sign translated."
[0152] 2. Terminal: The microphone captures the audio in MP3 format, and the camera takes a video of the sign in JPEG format. These data are sent to the server.
[0153] 3. Server: Uses a speech recognition engine (Google Speech-to-Text) to convert the voice data into text data such as "Please translate the text on the sign." OpenCV is used to extract the text area from the video of the sign, and the Google Translate API is used to translate it into the specified language.
[0154] 4. Server: Generates result data in JSON format, including the translation result "Business hours: 10:00-22:00", compresses it with GZIP, and sends it to the terminal.
[0155] 5. Terminal: The received result data is decompressed, and the translation result is converted into voice data using a speech synthesis engine (Google Text-to-Speech), and the user is notified that "This sign says 'Business hours: 10:00-22:00'."
[0156] Anomaly detection example
[0157] 1. User: (Unconscious, collapsed)
[0158] 2. Terminal: The camera captures the user's abnormal state in MPEG format and sends the abnormal state data to the server.
[0159] 3. Server: Analyzes abnormal situations using FaceNet and OpenPose and determines whether the user is in danger.
[0160] 4. Server: Contacts emergency services and provides the user's location and video data, and also notifies the user's emergency contacts.
[0161] 5. Terminal: The user is notified that an ambulance has been dispatched using a speech synthesis engine (Google Text-to-Speech).
[0162] Example prompts to input to the generative AI model
[0163] 1. For a translation request: "Please explain how the system works when a user wants the text on a sign translated."
[0164] 2. For anomaly detection: "If the user loses consciousness, explain how the system will detect the anomaly and respond."
[0165] In this way, the system of the present invention can provide users with information quickly and accurately in their daily lives and can also respond to emergencies.
[0166] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0167] Step 1:
[0168] The user speaks to the smart earphones. To request a translation, the user says, "I want the text on the sign translated." Voice data is generated as input.
[0169] Step 2:
[0170] The device uses a microphone to capture sound and generate digital audio data, which is then saved in WAV or MP3 format.
[0171] Step 3:
[0172] The device's first-person camera captures the user's viewpoint and generates video data. Specifically, high-resolution video data in JPEG or MPEG format is generated. The viewpoint video data is generated as input.
[0173] Step 4:
[0174] The audio and video data generated by the device is temporarily stored in local storage and compressed. Specifically, the data size is reduced using GZIP compression. The generated audio and video data is used as input, and a compressed data file is output.
[0175] Step 5:
[0176] The device sends the compressed audio and video data to a server over the Internet. Specifically, the data is sent securely using the HTTPS protocol. The compressed data file is used as input, and a response indicating that it has been sent to the server is output.
[0177] Step 6:
[0178] The server analyzes the received voice data and recognizes the user's intent. Specifically, it converts the voice data into text data using Google Speech-to-Text. The voice data is used as input, and the text data "I would like the words on the sign translated" is output.
[0179] Step 7:
[0180] The server recognizes from the text data that "I want the text on the sign translated" and analyzes the video data. Specifically, it uses OpenCV to extract the text area on the sign from the video data. The video data is used as input, and the coordinate information of the text area is output.
[0181] Step 8:
[0182] The server translates the extracted text into the specified language. Specifically, it uses the Google Translate API to translate the text. The text is used as input, and the translated text "Business hours: 10:00-22:00" is output.
[0183] Step 9:
[0184] The server generates result data containing the translation results and formats it in JSON format. The data is then compressed again using GZIP and sent to the terminal. Specifically, JSON-structured data is generated, compressed, and sent. The translation results are used as input, and a compressed result data file is output.
[0185] Step 10:
[0186] The device unpacks the received result data and uses Google Text-to-Speech to convert the translation results into audio data. Specifically, audio data such as "This sign says 'Business hours: 10:00-22:00'" is generated. The result data is used as input, and audio data is output.
[0187] Step 11:
[0188] The device notifies the user of the generated audio data, specifically by playing the audio through the smart earphones. The audio data is used as input and the user's listening is output.
[0189] (Application example 1)
[0190] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0191] In recent years, the introduction of smart devices and AI technology has progressed in many industries, and efficiency is also required in logistics centers. However, in the current work environment where a large amount of information and real-time instructions are required, employees have to manually check information and receive instructions, which is time-consuming. To solve this problem, a system is needed that allows workers to efficiently receive work instructions using audio and point-of-view video and respond immediately. Therefore, the present invention provides a smart device system that improves work efficiency in logistics centers.
[0192] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0193] In this invention, the server includes means for capturing user voice, means for capturing video from the user's viewpoint, means for transmitting the captured voice data and video data to the server, means for analyzing the voice data by the server to recognize the user's intention, means for analyzing the video data by the server to extract necessary information, means for generating a response based on the analyzed data and transmitting result data to the terminal, means for notifying the user of the transmitted result, and means for combining voice capture and video capture to provide information corresponding to a specific work instruction. This makes it possible for workers at a logistics center to issue voice instructions to provide a pick list, and for camera video to be analyzed to notify inventory information in real time.
[0194] Key Word Definitions
[0195] An "audio capture means" is a device or method that receives user-uttered audio and converts it into digital data.
[0196] A "means for capturing viewpoint images" is a device or method for acquiring images seen from the user's field of view in real time and recording them as digital data.
[0197] The "means for transmitting audio data and video data to a server" refers to a device or method for transferring the captured audio data and video data to a server via a communication network such as the Internet.
[0198] "Means for analyzing voice data and recognizing user intent" refers to algorithms or software that analyze received voice data and understand what the user is requesting.
[0199] "Means for analyzing video data and extracting necessary information" refers to algorithms or software that analyze received video data and identify and extract specific information (for example, text information or object information).
[0200] The "means for generating a response and transmitting result data to the terminal" refers to a device or method that generates an appropriate response based on the analyzed data and forwards the result to the user terminal.
[0201] "Means for notifying the user of the transmitted results" refers to a device or method for notifying the user of the transmitted results data at the terminal, including, for example, audio notification or visual display.
[0202] "Means for providing information corresponding to specific work instructions" refers to a method or device that combines audio and video capture to provide a user with specific instructions or information for a specific task or work.
[0203] A "pick list" is a list of instructions provided to workers for picking and sorting specific items at a distribution center.
[0204] "Inventory information" refers to data regarding the type, quantity, location, etc. of products stored in warehouses and logistics centers.
[0205] MODE FOR CARRYING OUT THE INVENTION
[0206] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system is particularly suitable for improving work efficiency in logistics centers. Specific embodiments for implementing this system are described below.
[0207] System configuration
[0208] This system consists of the following main components:
[0209] 1. User device (smart earphones and first-person camera)
[0210] 2. Server (AI processing unit)
[0211] 3. Internet connection
[0212] User terminal
[0213] The user terminal includes a microphone for capturing the user's voice and a first-person camera for capturing the user's point of view. The terminal also includes a communication module for transmitting data to a server.
[0214] server
[0215] The server includes an AI processing unit for analyzing the received audio and video data. This allows the server to recognize the user's intent and execute processing to provide the necessary information. The server also returns the resulting data to the device, which notifies the user of the results.
[0216] System Operation
[0217] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[0218] 1. Audio capture and transmission
[0219] The user speaks to the smart earphones, saying, "Tell me the next picklist." The device's microphone captures this voice and generates audio data. At the same time, a camera capturing the user's point of view captures the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed, and then sent to a server via the Internet.
[0220] 2. Analysis and processing on the server
[0221] The server analyzes the received voice data and recognizes the user's intent. For example, it extracts the intent "next pick list" from the voice. At the same time, it analyzes the video data and identifies shelf information, for example.
[0222] 3. Generating and returning result data
[0223] The server generates a picklist based on the analysis results, compresses this data again, and sends it to the terminal.
[0224] 4. Notice to Users
[0225] The terminal decompresses the received result data and notifies the user by voice, for example, by giving instructions such as "Please take product B from shelf A3."
[0226] Specific examples
[0227] Picklist Request Example
[0228] 1. User: "What's the next picklist?"
[0229] 2. Device: The microphone captures audio and the camera captures video of the surroundings. This data is sent to the server.
[0230] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the shelf information, then generates a pick list.
[0231] 4. Server: Sends the picklist to the terminal.
[0232] 5. Terminal: The picklist is announced to the user via voice: "Please pick up product B from shelf A3."
[0233] In this way, the system of the present invention can provide pick lists by voice instructions from logistics center workers and can analyze camera images to notify inventory information in real time. This system is expected to significantly improve work efficiency at logistics centers.
[0234] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0235] System program processing flow
[0236] Step 1:
[0237] The user speaks to the smart earphones and gives a voice command such as "Tell me the next picklist." This voice is captured by the microphone and generated as voice data.
[0238] Input: User's voice
[0239] Output: Audio data
[0240] How it works: The microphone in the smart earphones captures the user's voice and converts it into digital data.
[0241] Step 2:
[0242] The device's first-person camera captures images from the user's perspective and generates video data.
[0243] Input: Video from the user's point of view
[0244] Output: Video data
[0245] Specific operation: The device's camera captures video from the user's perspective in real time and stores it as digital data.
[0246] Step 3:
[0247] The captured audio and video data is temporarily stored in the terminal, compressed, and then transmitted to a server via the Internet.
[0248] Input: Audio data, video data
[0249] Output: Data sent to the server
[0250] Specific operation: Audio and video data is stored in the device's local storage, the data is compressed using a compression algorithm, and then sent to a server via an Internet connection.
[0251] Step 4:
[0252] The server analyzes the received voice data and recognizes the user's intent, for example, extracting the intent "next picklist."
[0253] Input: Audio data
[0254] Output: User intent (text format)
[0255] Specific operation: The server's AI processing unit uses a speech recognition engine (e.g., Google's speech recognition API) to convert the voice data into text and analyze the user's intent.
[0256] Step 5:
[0257] At the same time, the server analyzes the video data to identify, for example, shelf information.
[0258] Input: Video data
[0259] Output: Analysis results (shelf information)
[0260] Specific operation: Use the server's video analysis engine (such as OpenCV) to extract and identify specific information (e.g., shelf labels) from the video data.
[0261] Step 6:
[0262] The server generates a picklist based on the analysis results, compresses this data again, and sends it to the terminal.
[0263] Input: User intent, shelf information
[0264] Output: Picklist data
[0265] Specific operation: The generated picklist is compressed for transmission to the terminal, and then transmitted to the terminal via an Internet connection.
[0266] Step 7:
[0267] The terminal decompresses the received result data and notifies the user by voice, for example, by giving instructions such as "Please take product B from shelf A3."
[0268] Input: Compressed result data
[0269] Output: Audio notification
[0270] Specific operation: The device decompresses the compressed data and notifies the user by voice using a speech synthesis engine (e.g., Google TTS).
[0271] This series of processing steps allows workers at the logistics center to efficiently receive work instructions and respond immediately.
[0272] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0273] This invention is a smart earphone system that provides users with necessary information in real time in their daily lives and recognizes their emotions to provide appropriate feedback. This system captures audio and video from various viewpoints and analyzes the data on a server to generate a variety of responses. Furthermore, by combining it with an emotion engine, it is possible to provide adaptive services according to the user's emotions.
[0274] System configuration
[0275] This system consists of the following main components:
[0276] 1. User device (smart earphones and first-person camera)
[0277] 2. Server (AI processing unit and emotion engine)
[0278] 3. Internet connection
[0279] User terminal
[0280] The user terminal includes a microphone to capture the user's voice, a first-person camera to capture the user's point of view, and a communication module to transmit this data to a server via the Internet.
[0281] server
[0282] The server includes an AI processing unit for analyzing the received audio and video data. It is connected to an emotion engine, which is capable of recognizing the user's emotions from the received data. The server then generates appropriate information and responses based on the user's intentions and emotions, and notifies the user via the device.
[0283] System Operation
[0284] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[0285] Audio capture and transmission
[0286] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and converts it into digital audio data. At the same time, a first-person camera captures video from the user's point of view and converts it into digital video data. This data is temporarily stored in local storage, then compressed and sent to a server using a secure protocol.
[0287] Server parsing and processing
[0288] The server analyzes the received voice data and converts it into text data using a speech recognition engine. For example, it recognizes the intent of "I want it translated." At the same time, it analyzes the video data, identifies the text area on the sign, and extracts the text information. It then uses an emotion engine to recognize the user's emotions from the voice and video data, generating information to indicate the urgency of the request and provide adaptive feedback. The server then translates the text information into the specified language using a translation engine.
[0289] Generate and return result data
[0290] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[0291] User Notification
[0292] The device receives and decompresses the resulting data. It then notifies the user of the translation results and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[0293] Specific examples
[0294] Translation request example
[0295] 1. User: "I want the text on the sign translated."
[0296] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[0297] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the characters on the sign. The emotion engine analyzes the user's urgency and emotion from the voice and video. The translation engine then translates the characters into the specified language.
[0298] 4. Server: Sends translation results and emotion recognition results to the device.
[0299] 5. Terminal: The user is notified of the translation result and emotional feedback via voice, such as "This sign says 'Business hours: 10:00-22:00'. Has your problem been resolved?"
[0300] Anomaly detection example
[0301] 1. User: (Unconscious, collapsed)
[0302] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[0303] 3. Server: Analyzes the video data and anomaly detection data and determines whether the user is in danger.
[0304] 4. Server: Generates and sends data to contact emergency services and provide them with the necessary information, as well as to notify registered emergency contacts.
[0305] 5. Terminal: The user is notified by voice that the emergency call has been completed. The user is told, "An ambulance has been dispatched. Please wait a moment."
[0306] In this way, the system of the present invention can provide information quickly and accurately in the user's daily life and provide adaptive feedback according to the user's emotions.
[0307] The processing flow will be explained below.
[0308] Specific processing steps
[0309] Example 1: Processing a translation request
[0310] Step 1:
[0311] The user speaks to the smart earphones and says, "I want the text on the sign translated."
[0312] Step 2:
[0313] The device's microphone captures the user's voice, which is then converted into a digital format.
[0314] Step 3:
[0315] The device's first-person camera captures video from the user's point of view, and the video data is converted into a digital format.
[0316] Step 4:
[0317] The audio and video data captured by the device is temporarily stored in local storage.
[0318] Step 5:
[0319] The terminal transmits the audio and video data to the server using a secure protocol.
[0320] Step 6:
[0321] The server receives the voice data and converts it into text data using a voice recognition engine. For example, a command such as "Please translate the text on the sign" is converted into text information.
[0322] Step 7:
[0323] The server receives the video data and uses computer vision technology to identify text regions in the video, then extracts the identified text information.
[0324] Step 8:
[0325] The server combines the results of speech recognition and video analysis to understand the user's intent. In this case, it recognizes the intent to "translate the text on the sign."
[0326] Step 9:
[0327] The emotion engine analyzes audio and video data to recognize the user's emotions. For example, it determines that the user is feeling anxious.
[0328] Step 10:
[0329] The server translates the text information into the specified language using a translation engine, for example, translating a Japanese sign into English.
[0330] Step 11:
[0331] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[0332] Step 12:
[0333] The terminal receives and decompresses the resulting data.
[0334] Step 13:
[0335] The device will then provide the user with a voice message with the result data, for example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[0336] Example 2: Emergency call due to abnormality detection
[0337] Step 1:
[0338] The user falls into an abnormal state such as losing consciousness.
[0339] Step 2:
[0340] The device's first-person camera captures the user's unusual situation.
[0341] Step 3:
[0342] The device acquires anomaly detection data from the biometric sensor and temporarily stores it in local storage.
[0343] Step 4:
[0344] The video data and anomaly detection data captured by the device are sent to the server using a secure protocol.
[0345] Step 5:
[0346] The server analyzes the received video data and anomaly detection data.
[0347] Step 6:
[0348] The server determines from the analysis results that the user is in danger.
[0349] Step 7:
[0350] The emotion engine analyzes the data and determines that the user is in an emergency.
[0351] Step 8:
[0352] The server automatically contacts the nearest emergency services and provides them with the necessary information, such as the user's current location and condition.
[0353] Step 9:
[0354] The server generates data to notify the registered emergency contacts and transmits this to the terminal.
[0355] Step 10:
[0356] The device receives the emergency call results and provides the information to relatives and friends via voice or other notification means.
[0357] Step 11:
[0358] The terminal notifies the user by voice that the ambulance has been dispatched, for example, by saying, "The ambulance has been dispatched. Please wait a moment."
[0359] The above are the specific processing steps in the system of the present invention.
[0360] Example 2
[0361] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0362] Conventional information provision systems have had difficulty efficiently analyzing the user's voice and visual perspective and providing appropriate information and feedback in real time. They also lacked the technology to recognize the user's emotions and generate appropriate responses accordingly. Furthermore, there was no mechanism for detecting abnormal situations and responding quickly, which created problems in ensuring the user's safety.
[0363] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing voice data and converting it into text data using a voice recognition engine, means for analyzing video data and extracting information using a video analysis engine, and means for recognizing the user's emotions from the voice data and video data using an emotion engine. This makes it possible to analyze the user's voice and video from their viewpoint in real time and provide adaptive information and feedback according to the user's intentions and emotions. Furthermore, in an emergency, abnormalities can be detected and emergency services can be quickly contacted, thereby ensuring the user's safety.
[0364] "User" refers to an individual who uses the system and receives information.
[0365] "Means for capturing audio" refers to a microphone and its peripheral devices for converting a user's speech into digital data.
[0366] The term "means for capturing an image of a viewpoint" refers to a camera and its peripheral devices for converting an image seen from the user's viewpoint into digital data.
[0367] "Means for transmitting captured audio and video data to a server" refers to a communications module for transmitting digital data to a server via the Internet or other network.
[0368] "Means for analyzing voice data and converting it into text data using a voice recognition engine" refers to voice recognition technology for processing voice data and converting it into text information.
[0369] "Means for analyzing video data and extracting information using a video analytics engine" refers to video analytics technology for processing video data to extract specific information.
[0370] "Means for recognizing a user's emotions from audio data and video data using an emotion engine" refers to a technology for analyzing a user's emotions by understanding the characteristics of audio and video.
[0371] "Means for generating a response based on the analyzed data and transmitting the resulting data to the terminal" refers to algorithms and communication modules for generating an appropriate response based on the analyzed information and transmitting the response to the terminal.
[0372] "Means for notifying the user of the transmitted results" refers to a speaker or display on the terminal for displaying or notifying the user of the results by voice.
[0373] "Text data" refers to the text information of a user's speech converted by a voice recognition engine.
[0374] "Means for extracting information" refers to the processes and algorithms used by the video analytics engine to extract specific information.
[0375] "Means for recognizing emotions" refers to a process for analyzing and identifying a user's emotions using an emotion engine.
[0376] "Results data" refers to data that includes response information generated by a server.
[0377] This invention relates to a system that provides users with necessary information in real time in their daily lives and recognizes their emotions to provide appropriate feedback. This system captures audio and video from various viewpoints and analyzes the data on a computer server, enabling it to generate a variety of responses. Furthermore, by combining it with an emotion engine, it becomes possible to provide adaptive services according to the user's emotions.
[0378] System configuration
[0379] This system consists of the following main components:
[0380] 1. User device (smart earphones and first-person camera)
[0381] 2. Server (AI processing unit and emotion engine)
[0382] 3. Internet connection
[0383] User terminal
[0384] The user terminal includes a microphone to capture the user's voice, a first-person camera to capture the user's point of view, and a communication module to transmit this data to a server via the Internet.
[0385] server
[0386] The server includes an AI processing unit for analyzing the received audio and video data. It is connected to an emotion engine, which is capable of recognizing the user's emotions from the received data. The server then generates appropriate information and responses based on the user's intentions and emotions, and notifies the user via the device.
[0387] Specific operation example
[0388] Audio capture and transmission
[0389] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and converts it into digital audio data. At the same time, a first-person camera captures video from the user's point of view and converts it into digital video data. This data is temporarily stored in local storage, then compressed and sent to a server using a secure protocol.
[0390] Server parsing and processing
[0391] The server analyzes the received voice data and converts it into text data using a speech recognition engine. For example, it recognizes the intent of "I want it translated." At the same time, it analyzes the video data, identifies the text area on the sign, and extracts the text information. It then uses an emotion engine to recognize the user's emotions from the voice and video data, generating information to indicate the urgency of the request and provide adaptive feedback. The server then translates the text information into the specified language using a translation engine.
[0392] Generate and return result data
[0393] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[0394] User Notification
[0395] The device receives and decompresses the resulting data. It then notifies the user of the translation results and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[0396] Specific examples
[0397] Translation request example
[0398] 1. User: "I want the text on the sign translated."
[0399] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[0400] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the characters on the sign. The emotion engine analyzes the user's urgency and emotion from the voice and video. The translation engine then translates the characters into the specified language.
[0401] 4. Server: Sends translation results and emotion recognition results to the device.
[0402] 5. Terminal: The user is notified of the translation result and emotional feedback via voice, such as "This sign says 'Business hours: 10:00-22:00'. Has your problem been resolved?"
[0403] Anomaly detection example
[0404] 1. User: (Unconscious, collapsed)
[0405] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[0406] 3. Server: Analyzes the video data and anomaly detection data and determines whether the user is in danger.
[0407] 4. Server: Generates and sends data to contact emergency services and provide them with the necessary information, as well as to notify registered emergency contacts.
[0408] 5. Terminal: The user is notified by voice that the emergency call has been completed. The user is told, "An ambulance has been dispatched. Please wait a moment."
[0409] In this way, the system of the present invention can provide information quickly and accurately in the user's daily life and provide adaptive feedback according to the user's emotions.
[0410] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0411] Step 1:
[0412] The user speaks to the smart earphones, saying, "I want the text on the sign to be translated." The user's voice becomes the input.
[0413] Step 2:
[0414] The device's microphone captures the user's voice and converts it from an analog audio signal into digital audio data, which is then output.
[0415] Step 3:
[0416] The device's first-person camera captures video from the user's point of view and converts the analog video signal into digital video data, which is then output.
[0417] Step 4:
[0418] The device temporarily stores digital audio and video data in local storage, which acts as a temporary buffer.
[0419] Step 5:
[0420] The device compresses the stored data (e.g., in gzip format) to make it easier to analyze, and the compressed data is output.
[0421] Step 6:
[0422] The device sends the compressed data to the server using a secure protocol (e.g., HTTPS), which then becomes the server's input.
[0423] Step 7:
[0424] The server analyzes the received voice data using a voice recognition engine (e.g., general voice recognition technology) and converts the digital voice data into text data, which is then output.
[0425] Step 8:
[0426] The server analyzes the text data and recognizes the user's intent (e.g., a request to "translate"). The recognition result becomes the input for the next step.
[0427] Step 9:
[0428] The server analyzes the received video data using a video analysis engine (e.g., general video recognition technology) to identify the text area on the sign. The identified text area becomes the output.
[0429] Step 10:
[0430] The server extracts text information from the specified text area, and this text information becomes the output.
[0431] Step 11:
[0432] The server uses an emotion engine (e.g., general emotion analysis technology) to recognize the user's emotion from the audio and video data. The emotion recognition result is the output.
[0433] Step 12:
[0434] The server generates a response based on the analysis results (user intent, text information, and emotion recognition results). The generated response becomes the output.
[0435] Step 13:
[0436] The server compresses the response data (e.g., in gzip format) and sends it to the terminal. The sent data becomes the terminal's input.
[0437] Step 14:
[0438] The device decompresses the received data and provides the translation results and feedback to the user in a voice message with a tone and content that adapts to the user's emotions, such as, "This sign says 'Business hours: 10:00-22:00'. Is there anything I can help you with?"
[0439] (Application example 2)
[0440] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0441] Conventional work support systems in factories often provide limited information and basic feedback, lacking adaptive support based on the worker's emotions and situation. They also lack early detection and rapid response to abnormal situations during work. This makes it difficult to improve work efficiency and ensure safety.
[0442] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0443] In this invention, the server includes a means for recognizing the user's emotions and generating an appropriate response based on the emotions, a means for transmitting work instructions and safety guidelines in real time within the factory, and a means for detecting abnormalities from received video data and automatically contacting emergency services. This enables adaptive support based on the emotions of workers, as well as early detection and rapid response to abnormal situations during work.
[0444] "User" refers to an individual who uses the system, particularly a worker who performs work within a factory.
[0445] "Audio data" refers to digital data that captures audio information uttered by a user.
[0446] "Visual data" refers to digital data that contains video information captured from a user's point of view.
[0447] "Server" refers to a computing device that analyzes the transmitted audio and visual data, extracts the necessary information, and generates a response for the user.
[0448] "Emotion recognition" refers to the process of analyzing a user's audio and visual data to identify the user's emotional state.
[0449] "Work instructions" refers to instruction information that conveys specific work content and procedures to users within a factory.
[0450] "Safety guidelines" refers to instruction documents that clearly state the safety measures and precautions that should be observed during work.
[0451] "Anomaly detection" refers to the process of analyzing received video data to identify situations that deviate from normal working conditions or emergencies.
[0452] "Emergency services" refers to external support services such as ambulances, fire departments, and police departments that respond to emergency situations during work.
[0453] This invention is a smart earphone system that provides real-time support to workers (users) as a work support system in factories. The system captures and analyzes the user's voice and visual data to recognize emotions and provide adaptive work instructions and safety guidelines. It also has the ability to detect abnormalities and automatically contact emergency services.
[0454] System configuration
[0455] This system consists of the following main components:
[0456] 1. User Device
[0457] The user terminal includes a smart earphone and a first-person camera worn by the worker. The terminal is equipped with a microphone for capturing the user's voice, a camera for capturing first-person perspective video, and a communication module for transmitting the data to a server.
[0458] 2. Server
[0459] The server is equipped with an AI processing unit and emotion engine for receiving and analyzing voice and visual data. It has the ability to analyze voice data to recognize the user's intentions and visual data to extract necessary information. The emotion engine also recognizes the user's emotions and generates adaptive responses according to those emotions.
[0460] System Operation
[0461] When a user speaks to the smart earbuds, the system works as follows:
[0462] 1. Audio and visual data capture
[0463] When a user requests a work instruction, the smart earphones' microphone captures audio and the camera captures video from the user's point of view, and these digital data are temporarily stored in local storage before being transmitted to a server using a secure protocol.
[0464] 2. Analysis and processing on the server
[0465] The server analyzes the received voice data and converts it into text data using a speech recognition engine. At the same time, it analyzes the visual data to extract the necessary information. For example, if a user requests, "Please tell me the next step in the process," the server recognizes the user's intention and generates information about the next step.
[0466] 3. Emotion Recognition and Response Generation
[0467] The emotion engine recognizes the user's emotions from audio and visual data, generating information to determine the urgency of the request and provide adaptive feedback. For example, if the user is feeling stressed, it will provide a safety warning.
[0468] 4. Generating and returning result data
[0469] The server generates result data including analysis results, emotion recognition results, and work instructions, compresses it, and sends it to the terminal.
[0470] 5. Notice to Users
[0471] The device receives and decompresses the resulting data. It then provides the user with work instructions and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "Next, install part A. Safety first."
[0472] Specific examples
[0473] Work instruction example
[0474] The user speaks to the system, saying, "Please tell me the next steps," and the first-person camera captures the work environment. The system analyzes the audio and video, generates appropriate work instructions, and notifies the user by voice.
[0475] Example prompts to input to a generative AI model:
[0476] The user speaks to the smart earphones, "What are the next steps?", and the first-person camera captures the work environment. What steps can be taken to analyze the audio and video, generate appropriate work instructions, and provide voice prompts?
[0477] In this way, the system of the present invention can provide information quickly and accurately to workers in their daily work, and can provide adaptive feedback according to the user's emotions.
[0478] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0479] Step 1:
[0480] This is the step where the user speaks instructions to the smart earphone. For example, the user might say, "Please tell me the next step." This voice input triggers the start of the next processing step.
[0481] Step 2:
[0482] This is the step where the device captures the audio and converts it into digital audio data.
[0483] Input: User's voice
[0484] Output: Digital audio data
[0485] The device's microphone captures the user's voice, and an internal processor converts this voice into digital audio data.
[0486] Step 3:
[0487] The terminal simultaneously captures the visual data and converts it into digital video data.
[0488] Input: Video from the user's point of view
[0489] Output: Digital video data
[0490] A first-person camera mounted on the device captures video from the user's point of view and converts this video into digital data.
[0491] Step 4:
[0492] This is a step in which the terminal transmits the captured audio and video data to the server.
[0493] Input: Digital audio data, digital video data
[0494] Output: Send data to the server
[0495] A communication module in the terminal transmits the captured audio and video data to a server over a secure protocol.
[0496] Step 5:
[0497] This is the step where the server analyzes the received voice data and recognizes the user's intention.
[0498] Input: Digital audio data
[0499] Output: Text data and intent recognition results
[0500] The server's speech recognition engine converts the voice data into text data, analyzes it, and identifies the user's intent.
[0501] Step 6:
[0502] This is the step where the server analyzes the received visual data and extracts the necessary information.
[0503] Input: Digital video data
[0504] Output: Extracted information
[0505] The visual analytics engine analyzes the video data and identifies necessary information (such as the state of the work environment or specific objects).
[0506] Step 7:
[0507] This is the step where the server recognizes the user's emotions from the audio data and visual data.
[0508] Input: Digital audio data, digital video data
[0509] Output: User's emotional state
[0510] The emotion engine analyzes audio and video to identify the user's emotional state.
[0511] Step 8:
[0512] This is the step in which the server generates a response based on the analyzed data, creates result data, and sends it to the terminal.
[0513] Input: Intention recognition results, extracted information, emotional state
[0514] Output: Result data
[0515] The server generates an appropriate response based on the user's intention, information about the working environment, and emotional state, and creates result data, which is then compressed and sent to the terminal.
[0516] Step 9:
[0517] This is the step in which the terminal decompresses the received result data and notifies the user.
[0518] Input: Compressed result data
[0519] Output: Audio notification
[0520] The terminal decompresses the received result data and notifies the user by voice of appropriate work instructions and precautions, such as "Next, please install part A. Safety first."
[0521] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0522] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0523] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0524] [Second embodiment]
[0525] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0526] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0527] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0528] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0529] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0530] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0531] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0532] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0533] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0534] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0535] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0536] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0537] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system has a series of functions: capturing audio, capturing viewpoint video, analyzing data, generating results, and notifying the user of those results. Specific embodiments of the system are described below.
[0538] System configuration
[0539] This system consists of the following main components:
[0540] 1. User device (smart earphones and first-person camera)
[0541] 2. Server (AI processing unit)
[0542] 3. Internet connection
[0543] User terminal
[0544] The user terminal includes a microphone for capturing the user's voice and a first-person camera for capturing the user's point of view. The terminal also includes a communication module for transmitting data to a server.
[0545] server
[0546] The server includes an AI processing unit for analyzing the received audio and video data. This allows the server to recognize the user's intent and execute processing to provide the necessary information. The server also returns the resulting data to the device, which notifies the user of the results.
[0547] System Operation
[0548] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[0549] Audio capture and transmission
[0550] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and generates audio data. At the same time, a camera capturing the user's point of view captures video of the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed, and then sent to a server via the Internet.
[0551] Server parsing and processing
[0552] The server analyzes the received voice data and recognizes the user's intent. For example, it extracts the intent "signboard" and "please translate" from the voice. At the same time, it analyzes the video data and identifies the text area on the signboard. The identified text information is translated into the specified language by the translation engine.
[0553] Generate and return result data
[0554] The server generates result data including the translation result, compresses this data again, and transmits it to the terminal.
[0555] User Notification
[0556] The device decompresses the received result data and notifies the user by voice, for example, "This sign says 'Business hours: 10:00-22:00'."
[0557] Specific examples
[0558] Translation request example
[0559] 1. User: "I want the text on the sign translated."
[0560] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[0561] 3. Server: The speech recognition engine converts the voice data into text, the video analysis engine recognizes the characters on the sign, and the translation engine then translates the characters into the specified language.
[0562] 4. Server: Sends the translation results to the device.
[0563] 5. Terminal: The translation result is notified to the user by voice.
[0564] Anomaly detection example
[0565] 1. User: (Unconscious, collapsed)
[0566] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[0567] 3. Server: Analyzes abnormal situations and determines that the user is in danger.
[0568] 4. Server: Contacts emergency services and provides them with the necessary information, and also notifies emergency contacts.
[0569] 5. Terminal: The user is notified by voice that an ambulance has been dispatched.
[0570] In this way, the system of the present invention can provide users with information quickly and accurately in their daily lives and can also respond to emergencies.
[0571] The processing flow will be explained below.
[0572] Specific processing steps of the program
[0573] Example 1: Processing a translation request
[0574] Step 1:
[0575] The user speaks to the smart earphones and says, "I want the text on the sign translated."
[0576] Step 2:
[0577] The device's microphone captures the user's voice, which is then converted into digital data.
[0578] Step 3:
[0579] The device's camera captures an image of the sign from the user's point of view, and the image data is also converted into digital data.
[0580] Step 4:
[0581] The audio and video data captured by the device is temporarily stored in local storage.
[0582] Step 5:
[0583] The terminal transmits audio and video data to the server using a secure protocol.
[0584] Step 6:
[0585] The server receives the voice data and uses a speech recognition engine to convert the digital speech into text, for example, recognizing a command such as "Please translate the text on the sign."
[0586] Step 7:
[0587] The server receives the video data and uses computer vision techniques to identify text areas in the video, for example, to extract text on a sign.
[0588] Step 8:
[0589] The server combines the results of voice recognition and video analysis to understand the user's intent (in this case, "translating the text on the sign").
[0590] Step 9:
[0591] The server uses a translation engine to translate the sign text into the specified language, for example, translating Japanese text into English.
[0592] Step 10:
[0593] The server generates result data including the translation result, compresses it, and transmits it to the terminal.
[0594] Step 11:
[0595] The terminal receives and decompresses the resulting data.
[0596] Step 12:
[0597] The device will then notify the user of the translation results by voice, for example, saying, "This sign says 'Business hours: 10:00-22:00'."
[0598] Example 2: Emergency call due to abnormality detection
[0599] Step 1:
[0600] The user falls into an abnormal state such as losing consciousness.
[0601] Step 2:
[0602] The device's camera captures the user's unusual postures and situations.
[0603] Step 3:
[0604] The device acquires anomaly detection data from biometric sensors and other devices and temporarily stores it in local storage.
[0605] Step 4:
[0606] The video data and anomaly detection data captured by the device are sent to the server using a secure protocol.
[0607] Step 5:
[0608] The server analyzes the received video data and anomaly detection data.
[0609] Step 6:
[0610] The server determines from the analysis results that the user is in danger.
[0611] Step 7:
[0612] The server automatically contacts the nearest emergency services and provides them with the necessary information, such as the user's current location and condition.
[0613] Step 8:
[0614] The server generates data to notify the registered emergency contacts and transmits this to the terminal.
[0615] Step 9:
[0616] The device receives the emergency call results and provides the information to relatives and friends via voice or other notification means.
[0617] Step 10:
[0618] The terminal notifies the user by voice that the ambulance has been dispatched, for example, by saying, "The ambulance has been dispatched. Please wait a moment."
[0619] The above are the specific processing steps in the system of the present invention.
[0620] Example 1
[0621] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0622] Current speech recognition and translation systems can sometimes make it difficult for users to quickly and accurately obtain the information they need. Furthermore, there is a lack of systems that can monitor users' status in real time and provide appropriate responses in emergencies. Therefore, there is a need for systems that can provide users with quick and accurate information in everyday life and respond appropriately in emergencies.
[0623] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0624] In this invention, the server includes means for analyzing voice data using a voice recognition engine to recognize the user's intention, means for analyzing video data using image analysis means to extract necessary information, and means for translating the extracted information into a specified language using a translation engine, thereby making it possible to recognize the user's intention from the voice data, extract necessary information from the video data, and further translate it.
[0625] An "audio capturing means" is a device or software for collecting and digitizing a user's voice.
[0626] The "means for capturing viewpoint images" is a device or software for collecting images based on the user's viewpoint.
[0627] A "compression means" is a device or software that compresses audio and video data to reduce its size.
[0628] The "means for transmitting to the server" is a device or software for transmitting compressed audio data and video data to the server via the Internet.
[0629] "Means for analyzing and recognizing the user's intent using a voice recognition engine" refers to software that analyzes voice data on a server, converts it into text data, and understands the user's requests and intent.
[0630] "Means for analyzing and extracting necessary information using image analysis means" refers to software that analyzes video data on a server and identifies and extracts important information from it.
[0631] The "means for translating into a specified language using a translation engine" is software for translating extracted information into another language.
[0632] The "means for generating a response and transmitting the resulting data to the terminal" refers to software and communication devices for generating a response based on the analyzed and translated data and transmitting the resulting data to the terminal.
[0633] The "means for notifying the user using a voice synthesis engine" refers to software and a playback device for converting the result data into voice data and notifying the user.
[0634] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system has a series of functions: capturing audio, capturing viewpoint video, analyzing data, generating results, and notifying the user of those results. Specific embodiments of the system are described below.
[0635] System configuration
[0636] This system consists of the following main components:
[0637] 1. User device (smart earphones and first-person camera)
[0638] 2. Server (AI processing unit)
[0639] 3. Internet connection
[0640] User terminal
[0641] The user terminal includes a microphone to capture the user's voice and a first-person camera to capture the user's point of view. The terminal also has a communications module for transmitting data to a server. A digital microphone is used to capture the voice, and a high-resolution camera is used to capture the point of view. Data compression is performed using MP3 format for audio data and JPEG or MPEG format for video data.
[0642] server
[0643] The server includes an AI processing unit for analyzing the received audio and video data. For example, Google Speech-to-Text is used as the speech recognition engine, and OpenCV is used for image analysis. Furthermore, the Google Translate API is used as the translation engine. The analyzed and translated data is stored in JSON format and sent back to the device after GZIP compression. The device then uses a speech synthesis engine (e.g., Google Text-to-Speech) to notify the user.
[0644] System Operation
[0645] In an embodiment of the system, when a user speaks to the smart earphones, the system operates as follows: When the user speaks to the smart earphones, saying, "I want the words on the sign translated," the digital microphone on the user device captures this voice and generates audio data. At the same time, the first-person camera captures video of the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed into MP3 and JPEG formats, and then transmitted to a server via the Internet.
[0646] The server analyzes the received voice data using Google Speech-to-Text to recognize the user's intent. For example, it extracts the intents "signboard" and "translate" from the voice. At the same time, it analyzes the video data using OpenCV to identify the text area on the signboard. The identified text information is translated into the specified language using the Google Translate API.
[0647] The server generates result data containing the translation results, composes this data in JSON format, compresses it with GZIP, and sends it to the device. The device then decompresses the received result data and uses Google Text-to-Speech to notify the user aloud. For example, "This sign says 'Business hours: 10:00-22:00'."
[0648] Specific examples
[0649] Translation request example
[0650] 1. User: "I want the text on the sign translated."
[0651] 2. Terminal: The microphone captures the audio in MP3 format, and the camera takes a video of the sign in JPEG format. These data are sent to the server.
[0652] 3. Server: Uses a speech recognition engine (Google Speech-to-Text) to convert the voice data into text data such as "Please translate the text on the sign." OpenCV is used to extract the text area from the video of the sign, and the Google Translate API is used to translate it into the specified language.
[0653] 4. Server: Generates result data in JSON format, including the translation result "Business hours: 10:00-22:00", compresses it with GZIP, and sends it to the terminal.
[0654] 5. Terminal: The received result data is decompressed, and the translation result is converted into voice data using a speech synthesis engine (Google Text-to-Speech), and the user is notified that "This sign says 'Business hours: 10:00-22:00'."
[0655] Anomaly detection example
[0656] 1. User: (Unconscious, collapsed)
[0657] 2. Terminal: The camera captures the user's abnormal state in MPEG format and sends the abnormal state data to the server.
[0658] 3. Server: Analyzes abnormal situations using FaceNet and OpenPose and determines whether the user is in danger.
[0659] 4. Server: Contacts emergency services and provides the user's location and video data, and also notifies the user's emergency contacts.
[0660] 5. Terminal: The user is notified that an ambulance has been dispatched using a speech synthesis engine (Google Text-to-Speech).
[0661] Example prompts to input to the generative AI model
[0662] 1. For a translation request: "Please explain how the system works when a user wants the text on a sign translated."
[0663] 2. For anomaly detection: "If the user loses consciousness, explain how the system will detect the anomaly and respond."
[0664] In this way, the system of the present invention can provide users with information quickly and accurately in their daily lives and can also respond to emergencies.
[0665] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0666] Step 1:
[0667] The user speaks to the smart earphones. To request a translation, the user says, "I want the text on the sign translated." Voice data is generated as input.
[0668] Step 2:
[0669] The device uses a microphone to capture sound and generate digital audio data, which is then saved in WAV or MP3 format.
[0670] Step 3:
[0671] The device's first-person camera captures the user's viewpoint and generates video data. Specifically, high-resolution video data in JPEG or MPEG format is generated. The viewpoint video data is generated as input.
[0672] Step 4:
[0673] The audio and video data generated by the device is temporarily stored in local storage and compressed. Specifically, the data size is reduced using GZIP compression. The generated audio and video data is used as input, and a compressed data file is output.
[0674] Step 5:
[0675] The device sends the compressed audio and video data to a server over the Internet. Specifically, the data is sent securely using the HTTPS protocol. The compressed data file is used as input, and a response indicating that it has been sent to the server is output.
[0676] Step 6:
[0677] The server analyzes the received voice data and recognizes the user's intent. Specifically, it converts the voice data into text data using Google Speech-to-Text. The voice data is used as input, and the text data "I would like the words on the sign translated" is output.
[0678] Step 7:
[0679] The server recognizes from the text data that "I want the text on the sign translated" and analyzes the video data. Specifically, it uses OpenCV to extract the text area on the sign from the video data. The video data is used as input, and the coordinate information of the text area is output.
[0680] Step 8:
[0681] The server translates the extracted text into the specified language. Specifically, it uses the Google Translate API to translate the text. The text is used as input, and the translated text "Business hours: 10:00-22:00" is output.
[0682] Step 9:
[0683] The server generates result data containing the translation results and formats it in JSON format. The data is then compressed again using GZIP and sent to the terminal. Specifically, JSON-structured data is generated, compressed, and sent. The translation results are used as input, and a compressed result data file is output.
[0684] Step 10:
[0685] The device unpacks the received result data and uses Google Text-to-Speech to convert the translation results into audio data. Specifically, audio data such as "This sign says 'Business hours: 10:00-22:00'" is generated. The result data is used as input, and audio data is output.
[0686] Step 11:
[0687] The device notifies the user of the generated audio data, specifically by playing the audio through the smart earphones. The audio data is used as input and the user's listening is output.
[0688] (Application example 1)
[0689] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0690] In recent years, the introduction of smart devices and AI technology has progressed in many industries, and efficiency is also required in logistics centers. However, in the current work environment where a large amount of information and real-time instructions are required, employees have to manually check information and receive instructions, which is time-consuming. To solve this problem, a system is needed that allows workers to efficiently receive work instructions using audio and point-of-view video and respond immediately. Therefore, the present invention provides a smart device system that improves work efficiency in logistics centers.
[0691] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0692] In this invention, the server includes means for capturing user voice, means for capturing video from the user's viewpoint, means for transmitting the captured voice data and video data to the server, means for analyzing the voice data by the server to recognize the user's intention, means for analyzing the video data by the server to extract necessary information, means for generating a response based on the analyzed data and transmitting result data to the terminal, means for notifying the user of the transmitted result, and means for combining voice capture and video capture to provide information corresponding to a specific work instruction. This makes it possible for workers at a logistics center to issue voice instructions to provide a pick list, and for camera video to be analyzed to notify inventory information in real time.
[0693] Key Word Definitions
[0694] An "audio capture means" is a device or method that receives user-uttered audio and converts it into digital data.
[0695] A "means for capturing viewpoint images" is a device or method for acquiring images seen from the user's field of view in real time and recording them as digital data.
[0696] The "means for transmitting audio data and video data to a server" refers to a device or method for transferring the captured audio data and video data to a server via a communication network such as the Internet.
[0697] "Means for analyzing voice data and recognizing user intent" refers to algorithms or software that analyze received voice data and understand what the user is requesting.
[0698] "Means for analyzing video data and extracting necessary information" refers to algorithms or software that analyze received video data and identify and extract specific information (for example, text information or object information).
[0699] The "means for generating a response and transmitting result data to the terminal" refers to a device or method that generates an appropriate response based on the analyzed data and forwards the result to the user terminal.
[0700] "Means for notifying the user of the transmitted results" refers to a device or method for notifying the user of the transmitted results data at the terminal, including, for example, audio notification or visual display.
[0701] "Means for providing information corresponding to specific work instructions" refers to a method or device that combines audio and video capture to provide a user with specific instructions or information for a specific task or work.
[0702] A "pick list" is a list of instructions provided to workers for picking and sorting specific items at a distribution center.
[0703] "Inventory information" refers to data regarding the type, quantity, location, etc. of products stored in warehouses and logistics centers.
[0704] MODE FOR CARRYING OUT THE INVENTION
[0705] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system is particularly suitable for improving work efficiency in logistics centers. Specific embodiments for implementing this system are described below.
[0706] System configuration
[0707] This system consists of the following main components:
[0708] 1. User device (smart earphones and first-person camera)
[0709] 2. Server (AI processing unit)
[0710] 3. Internet connection
[0711] User terminal
[0712] The user terminal includes a microphone for capturing the user's voice and a first-person camera for capturing the user's point of view. The terminal also includes a communication module for transmitting data to a server.
[0713] server
[0714] The server includes an AI processing unit for analyzing the received audio and video data. This allows the server to recognize the user's intent and execute processing to provide the necessary information. The server also returns the resulting data to the device, which notifies the user of the results.
[0715] System Operation
[0716] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[0717] 1. Audio capture and transmission
[0718] The user speaks to the smart earphones, saying, "Tell me the next picklist." The device's microphone captures this voice and generates audio data. At the same time, a camera capturing the user's point of view captures the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed, and then sent to a server via the Internet.
[0719] 2. Analysis and processing on the server
[0720] The server analyzes the received voice data and recognizes the user's intent. For example, it extracts the intent "next pick list" from the voice. At the same time, it analyzes the video data and identifies shelf information, for example.
[0721] 3. Generating and returning result data
[0722] The server generates a picklist based on the analysis results, compresses this data again, and sends it to the terminal.
[0723] 4. Notice to Users
[0724] The terminal decompresses the received result data and notifies the user by voice, for example, by giving instructions such as "Please take product B from shelf A3."
[0725] Specific examples
[0726] Picklist Request Example
[0727] 1. User: "What's the next picklist?"
[0728] 2. Device: The microphone captures audio and the camera captures video of the surroundings. This data is sent to the server.
[0729] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the shelf information, then generates a pick list.
[0730] 4. Server: Sends the picklist to the terminal.
[0731] 5. Terminal: The picklist is announced to the user via voice: "Please pick up product B from shelf A3."
[0732] In this way, the system of the present invention can provide pick lists by voice instructions from logistics center workers and can analyze camera images to notify inventory information in real time. This system is expected to significantly improve work efficiency at logistics centers.
[0733] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0734] System program processing flow
[0735] Step 1:
[0736] The user speaks to the smart earphones and gives a voice command such as "Tell me the next picklist." This voice is captured by the microphone and generated as voice data.
[0737] Input: User's voice
[0738] Output: Audio data
[0739] How it works: The microphone in the smart earphones captures the user's voice and converts it into digital data.
[0740] Step 2:
[0741] The device's first-person camera captures images from the user's perspective and generates video data.
[0742] Input: Video from the user's point of view
[0743] Output: Video data
[0744] Specific operation: The device's camera captures video from the user's perspective in real time and stores it as digital data.
[0745] Step 3:
[0746] The captured audio and video data is temporarily stored in the terminal, compressed, and then transmitted to a server via the Internet.
[0747] Input: Audio data, video data
[0748] Output: Data sent to the server
[0749] Specific operation: Audio and video data is stored in the device's local storage, the data is compressed using a compression algorithm, and then sent to a server via an Internet connection.
[0750] Step 4:
[0751] The server analyzes the received voice data and recognizes the user's intent, for example, extracting the intent "next picklist."
[0752] Input: Audio data
[0753] Output: User intent (text format)
[0754] Specific operation: The server's AI processing unit uses a speech recognition engine (e.g., Google's speech recognition API) to convert the voice data into text and analyze the user's intent.
[0755] Step 5:
[0756] At the same time, the server analyzes the video data to identify, for example, shelf information.
[0757] Input: Video data
[0758] Output: Analysis results (shelf information)
[0759] Specific operation: Use the server's video analysis engine (such as OpenCV) to extract and identify specific information (e.g., shelf labels) from the video data.
[0760] Step 6:
[0761] The server generates a picklist based on the analysis results, compresses this data again, and sends it to the terminal.
[0762] Input: User intent, shelf information
[0763] Output: Picklist data
[0764] Specific operation: The generated picklist is compressed for transmission to the terminal, and then transmitted to the terminal via an Internet connection.
[0765] Step 7:
[0766] The terminal decompresses the received result data and notifies the user by voice, for example, by giving instructions such as "Please take product B from shelf A3."
[0767] Input: Compressed result data
[0768] Output: Audio notification
[0769] Specific operation: The device decompresses the compressed data and notifies the user by voice using a speech synthesis engine (e.g., Google TTS).
[0770] This series of processing steps allows workers at the logistics center to efficiently receive work instructions and respond immediately.
[0771] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0772] This invention is a smart earphone system that provides users with necessary information in real time in their daily lives and recognizes their emotions to provide appropriate feedback. This system captures audio and video from various viewpoints and analyzes the data on a server to generate a variety of responses. Furthermore, by combining it with an emotion engine, it is possible to provide adaptive services according to the user's emotions.
[0773] System configuration
[0774] This system consists of the following main components:
[0775] 1. User device (smart earphones and first-person camera)
[0776] 2. Server (AI processing unit and emotion engine)
[0777] 3. Internet connection
[0778] User terminal
[0779] The user terminal includes a microphone to capture the user's voice, a first-person camera to capture the user's point of view, and a communication module to transmit this data to a server via the Internet.
[0780] server
[0781] The server includes an AI processing unit for analyzing the received audio and video data. It is connected to an emotion engine, which is capable of recognizing the user's emotions from the received data. The server then generates appropriate information and responses based on the user's intentions and emotions, and notifies the user via the device.
[0782] System Operation
[0783] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[0784] Audio capture and transmission
[0785] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and converts it into digital audio data. At the same time, a first-person camera captures video from the user's point of view and converts it into digital video data. This data is temporarily stored in local storage, then compressed and sent to a server using a secure protocol.
[0786] Server parsing and processing
[0787] The server analyzes the received voice data and converts it into text data using a speech recognition engine. For example, it recognizes the intent of "I want it translated." At the same time, it analyzes the video data, identifies the text area on the sign, and extracts the text information. It then uses an emotion engine to recognize the user's emotions from the voice and video data, generating information to indicate the urgency of the request and provide adaptive feedback. The server then translates the text information into the specified language using a translation engine.
[0788] Generate and return result data
[0789] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[0790] User Notification
[0791] The device receives and decompresses the resulting data. It then notifies the user of the translation results and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[0792] Specific examples
[0793] Translation request example
[0794] 1. User: "I want the text on the sign translated."
[0795] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[0796] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the characters on the sign. The emotion engine analyzes the user's urgency and emotion from the voice and video. The translation engine then translates the characters into the specified language.
[0797] 4. Server: Sends translation results and emotion recognition results to the device.
[0798] 5. Terminal: The user is notified of the translation result and emotional feedback via voice, such as "This sign says 'Business hours: 10:00-22:00'. Has your problem been resolved?"
[0799] Anomaly detection example
[0800] 1. User: (Unconscious, collapsed)
[0801] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[0802] 3. Server: Analyzes the video data and anomaly detection data and determines whether the user is in danger.
[0803] 4. Server: Generates and sends data to contact emergency services and provide them with the necessary information, as well as to notify registered emergency contacts.
[0804] 5. Terminal: The user is notified by voice that the emergency call has been completed. The user is told, "An ambulance has been dispatched. Please wait a moment."
[0805] In this way, the system of the present invention can provide information quickly and accurately in the user's daily life and provide adaptive feedback according to the user's emotions.
[0806] The processing flow will be explained below.
[0807] Specific processing steps
[0808] Example 1: Processing a translation request
[0809] Step 1:
[0810] The user speaks to the smart earphones and says, "I want the text on the sign translated."
[0811] Step 2:
[0812] The device's microphone captures the user's voice, which is then converted into a digital format.
[0813] Step 3:
[0814] The device's first-person camera captures video from the user's point of view, and the video data is converted into a digital format.
[0815] Step 4:
[0816] The audio and video data captured by the device is temporarily stored in local storage.
[0817] Step 5:
[0818] The terminal transmits the audio and video data to the server using a secure protocol.
[0819] Step 6:
[0820] The server receives the voice data and converts it into text data using a voice recognition engine. For example, a command such as "Please translate the text on the sign" is converted into text information.
[0821] Step 7:
[0822] The server receives the video data and uses computer vision technology to identify text regions in the video, then extracts the identified text information.
[0823] Step 8:
[0824] The server combines the results of speech recognition and video analysis to understand the user's intent. In this case, it recognizes the intent to "translate the text on the sign."
[0825] Step 9:
[0826] The emotion engine analyzes audio and video data to recognize the user's emotions. For example, it determines that the user is feeling anxious.
[0827] Step 10:
[0828] The server translates the text information into the specified language using a translation engine, for example, translating a Japanese sign into English.
[0829] Step 11:
[0830] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[0831] Step 12:
[0832] The terminal receives and decompresses the resulting data.
[0833] Step 13:
[0834] The device will then provide the user with a voice message with the result data, for example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[0835] Example 2: Emergency call due to abnormality detection
[0836] Step 1:
[0837] The user falls into an abnormal state such as losing consciousness.
[0838] Step 2:
[0839] The device's first-person camera captures the user's unusual situation.
[0840] Step 3:
[0841] The device acquires anomaly detection data from the biometric sensor and temporarily stores it in local storage.
[0842] Step 4:
[0843] The video data and anomaly detection data captured by the device are sent to the server using a secure protocol.
[0844] Step 5:
[0845] The server analyzes the received video data and anomaly detection data.
[0846] Step 6:
[0847] The server determines from the analysis results that the user is in danger.
[0848] Step 7:
[0849] The emotion engine analyzes the data and determines that the user is in an emergency.
[0850] Step 8:
[0851] The server automatically contacts the nearest emergency services and provides them with the necessary information, such as the user's current location and condition.
[0852] Step 9:
[0853] The server generates data to notify the registered emergency contacts and transmits this to the terminal.
[0854] Step 10:
[0855] The device receives the emergency call results and provides the information to relatives and friends via voice or other notification means.
[0856] Step 11:
[0857] The terminal notifies the user by voice that the ambulance has been dispatched, for example, by saying, "The ambulance has been dispatched. Please wait a moment."
[0858] The above are the specific processing steps in the system of the present invention.
[0859] Example 2
[0860] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0861] Conventional information provision systems have had difficulty efficiently analyzing the user's voice and visual perspective and providing appropriate information and feedback in real time. They also lacked the technology to recognize the user's emotions and generate appropriate responses accordingly. Furthermore, there was no mechanism for detecting abnormal situations and responding quickly, which created problems in ensuring the user's safety.
[0862] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing voice data and converting it into text data using a voice recognition engine, means for analyzing video data and extracting information using a video analysis engine, and means for recognizing the user's emotions from the voice data and video data using an emotion engine. This makes it possible to analyze the user's voice and video from their viewpoint in real time and provide adaptive information and feedback according to the user's intentions and emotions. Furthermore, in an emergency, abnormalities can be detected and emergency services can be quickly contacted, thereby ensuring the user's safety.
[0863] "User" refers to an individual who uses the system and receives information.
[0864] "Means for capturing audio" refers to a microphone and its peripheral devices for converting a user's speech into digital data.
[0865] The term "means for capturing an image of a viewpoint" refers to a camera and its peripheral devices for converting an image seen from the user's viewpoint into digital data.
[0866] "Means for transmitting captured audio and video data to a server" refers to a communications module for transmitting digital data to a server via the Internet or other network.
[0867] "Means for analyzing voice data and converting it into text data using a voice recognition engine" refers to voice recognition technology for processing voice data and converting it into text information.
[0868] "Means for analyzing video data and extracting information using a video analytics engine" refers to video analytics technology for processing video data to extract specific information.
[0869] "Means for recognizing a user's emotions from audio data and video data using an emotion engine" refers to a technology for analyzing a user's emotions by understanding the characteristics of audio and video.
[0870] "Means for generating a response based on the analyzed data and transmitting the resulting data to the terminal" refers to algorithms and communication modules for generating an appropriate response based on the analyzed information and transmitting the response to the terminal.
[0871] "Means for notifying the user of the transmitted results" refers to a speaker or display on the terminal for displaying or notifying the user of the results by voice.
[0872] "Text data" refers to the text information of a user's speech converted by a voice recognition engine.
[0873] "Means for extracting information" refers to the processes and algorithms used by the video analytics engine to extract specific information.
[0874] "Means for recognizing emotions" refers to a process for analyzing and identifying a user's emotions using an emotion engine.
[0875] "Results data" refers to data that includes response information generated by a server.
[0876] This invention relates to a system that provides users with necessary information in real time in their daily lives and recognizes their emotions to provide appropriate feedback. This system captures audio and video from various viewpoints and analyzes the data on a computer server, enabling it to generate a variety of responses. Furthermore, by combining it with an emotion engine, it becomes possible to provide adaptive services according to the user's emotions.
[0877] System configuration
[0878] This system consists of the following main components:
[0879] 1. User device (smart earphones and first-person camera)
[0880] 2. Server (AI processing unit and emotion engine)
[0881] 3. Internet connection
[0882] User terminal
[0883] The user terminal includes a microphone to capture the user's voice, a first-person camera to capture the user's point of view, and a communication module to transmit this data to a server via the Internet.
[0884] server
[0885] The server includes an AI processing unit for analyzing the received audio and video data. It is connected to an emotion engine, which is capable of recognizing the user's emotions from the received data. The server then generates appropriate information and responses based on the user's intentions and emotions, and notifies the user via the device.
[0886] Specific operation example
[0887] Audio capture and transmission
[0888] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and converts it into digital audio data. At the same time, a first-person camera captures video from the user's point of view and converts it into digital video data. This data is temporarily stored in local storage, then compressed and sent to a server using a secure protocol.
[0889] Server parsing and processing
[0890] The server analyzes the received voice data and converts it into text data using a speech recognition engine. For example, it recognizes the intent of "I want it translated." At the same time, it analyzes the video data, identifies the text area on the sign, and extracts the text information. It then uses an emotion engine to recognize the user's emotions from the voice and video data, generating information to indicate the urgency of the request and provide adaptive feedback. The server then translates the text information into the specified language using a translation engine.
[0891] Generate and return result data
[0892] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[0893] User Notification
[0894] The device receives and decompresses the resulting data. It then notifies the user of the translation results and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[0895] Specific examples
[0896] Translation request example
[0897] 1. User: "I want the text on the sign translated."
[0898] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[0899] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the characters on the sign. The emotion engine analyzes the user's urgency and emotion from the voice and video. The translation engine then translates the characters into the specified language.
[0900] 4. Server: Sends translation results and emotion recognition results to the device.
[0901] 5. Terminal: The user is notified of the translation result and emotional feedback via voice, such as "This sign says 'Business hours: 10:00-22:00'. Has your problem been resolved?"
[0902] Anomaly detection example
[0903] 1. User: (Unconscious, collapsed)
[0904] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[0905] 3. Server: Analyzes the video data and anomaly detection data and determines whether the user is in danger.
[0906] 4. Server: Generates and sends data to contact emergency services and provide them with the necessary information, as well as to notify registered emergency contacts.
[0907] 5. Terminal: The user is notified by voice that the emergency call has been completed. The user is told, "An ambulance has been dispatched. Please wait a moment."
[0908] In this way, the system of the present invention can provide information quickly and accurately in the user's daily life and provide adaptive feedback according to the user's emotions.
[0909] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0910] Step 1:
[0911] The user speaks to the smart earphones, saying, "I want the text on the sign to be translated." The user's voice becomes the input.
[0912] Step 2:
[0913] The device's microphone captures the user's voice and converts it from an analog audio signal into digital audio data, which is then output.
[0914] Step 3:
[0915] The device's first-person camera captures video from the user's point of view and converts the analog video signal into digital video data, which is then output.
[0916] Step 4:
[0917] The device temporarily stores digital audio and video data in local storage, which acts as a temporary buffer.
[0918] Step 5:
[0919] The device compresses the stored data (e.g., in gzip format) to make it easier to analyze, and the compressed data is output.
[0920] Step 6:
[0921] The device sends the compressed data to the server using a secure protocol (e.g., HTTPS), which then becomes the server's input.
[0922] Step 7:
[0923] The server analyzes the received voice data using a voice recognition engine (e.g., general voice recognition technology) and converts the digital voice data into text data, which is then output.
[0924] Step 8:
[0925] The server analyzes the text data and recognizes the user's intent (e.g., a request to "translate"). The recognition result becomes the input for the next step.
[0926] Step 9:
[0927] The server analyzes the received video data using a video analysis engine (e.g., general video recognition technology) to identify the text area on the sign. The identified text area becomes the output.
[0928] Step 10:
[0929] The server extracts text information from the specified text area, and this text information becomes the output.
[0930] Step 11:
[0931] The server uses an emotion engine (e.g., general emotion analysis technology) to recognize the user's emotion from the audio and video data. The emotion recognition result is the output.
[0932] Step 12:
[0933] The server generates a response based on the analysis results (user intent, text information, and emotion recognition results). The generated response becomes the output.
[0934] Step 13:
[0935] The server compresses the response data (e.g., in gzip format) and sends it to the terminal. The sent data becomes the terminal's input.
[0936] Step 14:
[0937] The device decompresses the received data and provides the translation results and feedback to the user in a voice message with a tone and content that adapts to the user's emotions, such as, "This sign says 'Business hours: 10:00-22:00'. Is there anything I can help you with?"
[0938] (Application example 2)
[0939] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0940] Conventional work support systems in factories often provide limited information and basic feedback, lacking adaptive support based on the worker's emotions and situation. They also lack early detection and rapid response to abnormal situations during work. This makes it difficult to improve work efficiency and ensure safety.
[0941] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0942] In this invention, the server includes a means for recognizing the user's emotions and generating an appropriate response based on the emotions, a means for transmitting work instructions and safety guidelines in real time within the factory, and a means for detecting abnormalities from received video data and automatically contacting emergency services. This enables adaptive support based on the emotions of workers, as well as early detection and rapid response to abnormal situations during work.
[0943] "User" refers to an individual who uses the system, particularly a worker who performs work within a factory.
[0944] "Audio data" refers to digital data that captures audio information uttered by a user.
[0945] "Visual data" refers to digital data that contains video information captured from a user's point of view.
[0946] "Server" refers to a computing device that analyzes the transmitted audio and visual data, extracts the necessary information, and generates a response for the user.
[0947] "Emotion recognition" refers to the process of analyzing a user's audio and visual data to identify the user's emotional state.
[0948] "Work instructions" refers to instruction information that conveys specific work content and procedures to users within a factory.
[0949] "Safety guidelines" refers to instruction documents that clearly state the safety measures and precautions that should be observed during work.
[0950] "Anomaly detection" refers to the process of analyzing received video data to identify situations that deviate from normal working conditions or emergencies.
[0951] "Emergency services" refers to external support services such as ambulances, fire departments, and police departments that respond to emergency situations during work.
[0952] This invention is a smart earphone system that provides real-time support to workers (users) as a work support system in factories. The system captures and analyzes the user's voice and visual data to recognize emotions and provide adaptive work instructions and safety guidelines. It also has the ability to detect abnormalities and automatically contact emergency services.
[0953] System configuration
[0954] This system consists of the following main components:
[0955] 1. User Device
[0956] The user terminal includes a smart earphone and a first-person camera worn by the worker. The terminal is equipped with a microphone for capturing the user's voice, a camera for capturing first-person perspective video, and a communication module for transmitting the data to a server.
[0957] 2. Server
[0958] The server is equipped with an AI processing unit and emotion engine for receiving and analyzing voice and visual data. It has the ability to analyze voice data to recognize the user's intentions and visual data to extract necessary information. The emotion engine also recognizes the user's emotions and generates adaptive responses according to those emotions.
[0959] System Operation
[0960] When a user speaks to the smart earbuds, the system works as follows:
[0961] 1. Audio and visual data capture
[0962] When a user requests a work instruction, the smart earphones' microphone captures audio and the camera captures video from the user's point of view, and these digital data are temporarily stored in local storage before being transmitted to a server using a secure protocol.
[0963] 2. Analysis and processing on the server
[0964] The server analyzes the received voice data and converts it into text data using a speech recognition engine. At the same time, it analyzes the visual data to extract the necessary information. For example, if a user requests, "Please tell me the next step in the process," the server recognizes the user's intention and generates information about the next step.
[0965] 3. Emotion Recognition and Response Generation
[0966] The emotion engine recognizes the user's emotions from audio and visual data, generating information to determine the urgency of the request and provide adaptive feedback. For example, if the user is feeling stressed, it will provide a safety warning.
[0967] 4. Generating and returning result data
[0968] The server generates result data including analysis results, emotion recognition results, and work instructions, compresses it, and sends it to the terminal.
[0969] 5. Notice to Users
[0970] The device receives and decompresses the resulting data. It then provides the user with work instructions and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "Next, install part A. Safety first."
[0971] Specific examples
[0972] Work instruction example
[0973] The user speaks to the system, saying, "Please tell me the next steps," and the first-person camera captures the work environment. The system analyzes the audio and video, generates appropriate work instructions, and notifies the user by voice.
[0974] Example prompts to input to a generative AI model:
[0975] The user speaks to the smart earphones, "What are the next steps?", and the first-person camera captures the work environment. What steps can be taken to analyze the audio and video, generate appropriate work instructions, and provide voice prompts?
[0976] In this way, the system of the present invention can provide information quickly and accurately to workers in their daily work, and can provide adaptive feedback according to the user's emotions.
[0977] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0978] Step 1:
[0979] This is the step where the user speaks instructions to the smart earphone. For example, the user might say, "Please tell me the next step." This voice input triggers the start of the next processing step.
[0980] Step 2:
[0981] This is the step where the device captures the audio and converts it into digital audio data.
[0982] Input: User's voice
[0983] Output: Digital audio data
[0984] The device's microphone captures the user's voice, and an internal processor converts this voice into digital audio data.
[0985] Step 3:
[0986] The terminal simultaneously captures the visual data and converts it into digital video data.
[0987] Input: Video from the user's point of view
[0988] Output: Digital video data
[0989] A first-person camera mounted on the device captures video from the user's point of view and converts this video into digital data.
[0990] Step 4:
[0991] This is a step in which the terminal transmits the captured audio and video data to the server.
[0992] Input: Digital audio data, digital video data
[0993] Output: Send data to the server
[0994] A communication module in the terminal transmits the captured audio and video data to a server over a secure protocol.
[0995] Step 5:
[0996] This is the step where the server analyzes the received voice data and recognizes the user's intention.
[0997] Input: Digital audio data
[0998] Output: Text data and intent recognition results
[0999] The server's speech recognition engine converts the voice data into text data, analyzes it, and identifies the user's intent.
[1000] Step 6:
[1001] This is the step where the server analyzes the received visual data and extracts the necessary information.
[1002] Input: Digital video data
[1003] Output: Extracted information
[1004] The visual analytics engine analyzes the video data and identifies necessary information (such as the state of the work environment or specific objects).
[1005] Step 7:
[1006] This is the step where the server recognizes the user's emotions from the audio data and visual data.
[1007] Input: Digital audio data, digital video data
[1008] Output: User's emotional state
[1009] The emotion engine analyzes audio and video to identify the user's emotional state.
[1010] Step 8:
[1011] This is the step in which the server generates a response based on the analyzed data, creates result data, and sends it to the terminal.
[1012] Input: Intention recognition results, extracted information, emotional state
[1013] Output: Result data
[1014] The server generates an appropriate response based on the user's intention, information about the working environment, and emotional state, and creates result data, which is then compressed and sent to the terminal.
[1015] Step 9:
[1016] This is the step in which the terminal decompresses the received result data and notifies the user.
[1017] Input: Compressed result data
[1018] Output: Audio notification
[1019] The terminal decompresses the received result data and notifies the user by voice of appropriate work instructions and precautions, such as "Next, please install part A. Safety first."
[1020] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1021] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1022] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1023] [Third embodiment]
[1024] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1025] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1027] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1028] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1029] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1031] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1032] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1034] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1035] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1036] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system has a series of functions: capturing audio, capturing viewpoint video, analyzing data, generating results, and notifying the user of those results. Specific embodiments of the system are described below.
[1037] System configuration
[1038] This system consists of the following main components:
[1039] 1. User device (smart earphones and first-person camera)
[1040] 2. Server (AI processing unit)
[1041] 3. Internet connection
[1042] User terminal
[1043] The user terminal includes a microphone for capturing the user's voice and a first-person camera for capturing the user's point of view. The terminal also includes a communication module for transmitting data to a server.
[1044] server
[1045] The server includes an AI processing unit for analyzing the received audio and video data. This allows the server to recognize the user's intent and execute processing to provide the necessary information. The server also returns the resulting data to the device, which notifies the user of the results.
[1046] System Operation
[1047] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[1048] Audio capture and transmission
[1049] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and generates audio data. At the same time, a camera capturing the user's point of view captures video of the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed, and then sent to a server via the Internet.
[1050] Server parsing and processing
[1051] The server analyzes the received voice data and recognizes the user's intent. For example, it extracts the intent "signboard" and "please translate" from the voice. At the same time, it analyzes the video data and identifies the text area on the signboard. The identified text information is translated into the specified language by the translation engine.
[1052] Generate and return result data
[1053] The server generates result data including the translation result, compresses this data again, and transmits it to the terminal.
[1054] User Notification
[1055] The device decompresses the received result data and notifies the user by voice, for example, "This sign says 'Business hours: 10:00-22:00'."
[1056] Specific examples
[1057] Translation request example
[1058] 1. User: "I want the text on the sign translated."
[1059] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[1060] 3. Server: The speech recognition engine converts the voice data into text, the video analysis engine recognizes the characters on the sign, and the translation engine then translates the characters into the specified language.
[1061] 4. Server: Sends the translation results to the device.
[1062] 5. Terminal: The translation result is notified to the user by voice.
[1063] Anomaly detection example
[1064] 1. User: (Unconscious, collapsed)
[1065] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[1066] 3. Server: Analyzes abnormal situations and determines that the user is in danger.
[1067] 4. Server: Contacts emergency services and provides them with the necessary information, and also notifies emergency contacts.
[1068] 5. Terminal: The user is notified by voice that an ambulance has been dispatched.
[1069] In this way, the system of the present invention can provide users with information quickly and accurately in their daily lives and can also respond to emergencies.
[1070] The processing flow will be explained below.
[1071] Specific processing steps of the program
[1072] Example 1: Processing a translation request
[1073] Step 1:
[1074] The user speaks to the smart earphones and says, "I want the text on the sign translated."
[1075] Step 2:
[1076] The device's microphone captures the user's voice, which is then converted into digital data.
[1077] Step 3:
[1078] The device's camera captures an image of the sign from the user's point of view, and the image data is also converted into digital data.
[1079] Step 4:
[1080] The audio and video data captured by the device is temporarily stored in local storage.
[1081] Step 5:
[1082] The terminal transmits audio and video data to the server using a secure protocol.
[1083] Step 6:
[1084] The server receives the voice data and uses a speech recognition engine to convert the digital speech into text, for example, recognizing a command such as "Please translate the text on the sign."
[1085] Step 7:
[1086] The server receives the video data and uses computer vision techniques to identify text areas in the video, for example, to extract text on a sign.
[1087] Step 8:
[1088] The server combines the results of voice recognition and video analysis to understand the user's intent (in this case, "translating the text on the sign").
[1089] Step 9:
[1090] The server uses a translation engine to translate the sign text into the specified language, for example, translating Japanese text into English.
[1091] Step 10:
[1092] The server generates result data including the translation result, compresses it, and transmits it to the terminal.
[1093] Step 11:
[1094] The terminal receives and decompresses the resulting data.
[1095] Step 12:
[1096] The device will then notify the user of the translation results by voice, for example, saying, "This sign says 'Business hours: 10:00-22:00'."
[1097] Example 2: Emergency call due to abnormality detection
[1098] Step 1:
[1099] The user falls into an abnormal state such as losing consciousness.
[1100] Step 2:
[1101] The device's camera captures the user's unusual postures and situations.
[1102] Step 3:
[1103] The device acquires anomaly detection data from biometric sensors and other devices and temporarily stores it in local storage.
[1104] Step 4:
[1105] The video data and anomaly detection data captured by the device are sent to the server using a secure protocol.
[1106] Step 5:
[1107] The server analyzes the received video data and anomaly detection data.
[1108] Step 6:
[1109] The server determines from the analysis results that the user is in danger.
[1110] Step 7:
[1111] The server automatically contacts the nearest emergency services and provides them with the necessary information, such as the user's current location and condition.
[1112] Step 8:
[1113] The server generates data to notify the registered emergency contacts and transmits this to the terminal.
[1114] Step 9:
[1115] The device receives the emergency call results and provides the information to relatives and friends via voice or other notification means.
[1116] Step 10:
[1117] The terminal notifies the user by voice that the ambulance has been dispatched, for example, by saying, "The ambulance has been dispatched. Please wait a moment."
[1118] The above are the specific processing steps in the system of the present invention.
[1119] Example 1
[1120] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1121] Current speech recognition and translation systems can sometimes make it difficult for users to quickly and accurately obtain the information they need. Furthermore, there is a lack of systems that can monitor users' status in real time and provide appropriate responses in emergencies. Therefore, there is a need for systems that can provide users with quick and accurate information in everyday life and respond appropriately in emergencies.
[1122] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1123] In this invention, the server includes means for analyzing voice data using a voice recognition engine to recognize the user's intention, means for analyzing video data using image analysis means to extract necessary information, and means for translating the extracted information into a specified language using a translation engine, thereby making it possible to recognize the user's intention from the voice data, extract necessary information from the video data, and further translate it.
[1124] An "audio capturing means" is a device or software for collecting and digitizing a user's voice.
[1125] The "means for capturing viewpoint images" is a device or software for collecting images based on the user's viewpoint.
[1126] A "compression means" is a device or software that compresses audio and video data to reduce its size.
[1127] The "means for transmitting to the server" is a device or software for transmitting compressed audio data and video data to the server via the Internet.
[1128] "Means for analyzing and recognizing the user's intent using a voice recognition engine" refers to software that analyzes voice data on a server, converts it into text data, and understands the user's requests and intent.
[1129] "Means for analyzing and extracting necessary information using image analysis means" refers to software that analyzes video data on a server and identifies and extracts important information from it.
[1130] The "means for translating into a specified language using a translation engine" is software for translating extracted information into another language.
[1131] The "means for generating a response and transmitting the resulting data to the terminal" refers to software and communication devices for generating a response based on the analyzed and translated data and transmitting the resulting data to the terminal.
[1132] The "means for notifying the user using a voice synthesis engine" refers to software and a playback device for converting the result data into voice data and notifying the user.
[1133] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system has a series of functions: capturing audio, capturing viewpoint video, analyzing data, generating results, and notifying the user of those results. Specific embodiments of the system are described below.
[1134] System configuration
[1135] This system consists of the following main components:
[1136] 1. User device (smart earphones and first-person camera)
[1137] 2. Server (AI processing unit)
[1138] 3. Internet connection
[1139] User terminal
[1140] The user terminal includes a microphone to capture the user's voice and a first-person camera to capture the user's point of view. The terminal also has a communications module for transmitting data to a server. A digital microphone is used to capture the voice, and a high-resolution camera is used to capture the point of view. Data compression is performed using MP3 format for audio data and JPEG or MPEG format for video data.
[1141] server
[1142] The server includes an AI processing unit for analyzing the received audio and video data. For example, Google Speech-to-Text is used as the speech recognition engine, and OpenCV is used for image analysis. Furthermore, the Google Translate API is used as the translation engine. The analyzed and translated data is stored in JSON format and sent back to the device after GZIP compression. The device then uses a speech synthesis engine (e.g., Google Text-to-Speech) to notify the user.
[1143] System Operation
[1144] In an embodiment of the system, when a user speaks to the smart earphones, the system operates as follows: When the user speaks to the smart earphones, saying, "I want the words on the sign translated," the digital microphone on the user device captures this voice and generates audio data. At the same time, the first-person camera captures video of the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed into MP3 and JPEG formats, and then transmitted to a server via the Internet.
[1145] The server analyzes the received voice data using Google Speech-to-Text to recognize the user's intent. For example, it extracts the intents "signboard" and "translate" from the voice. At the same time, it analyzes the video data using OpenCV to identify the text area on the signboard. The identified text information is translated into the specified language using the Google Translate API.
[1146] The server generates result data containing the translation results, composes this data in JSON format, compresses it with GZIP, and sends it to the device. The device then decompresses the received result data and uses Google Text-to-Speech to notify the user aloud. For example, "This sign says 'Business hours: 10:00-22:00'."
[1147] Specific examples
[1148] Translation request example
[1149] 1. User: "I want the text on the sign translated."
[1150] 2. Terminal: The microphone captures the audio in MP3 format, and the camera takes a video of the sign in JPEG format. These data are sent to the server.
[1151] 3. Server: Uses a speech recognition engine (Google Speech-to-Text) to convert the voice data into text data such as "Please translate the text on the sign." OpenCV is used to extract the text area from the video of the sign, and the Google Translate API is used to translate it into the specified language.
[1152] 4. Server: Generates result data in JSON format, including the translation result "Business hours: 10:00-22:00", compresses it with GZIP, and sends it to the terminal.
[1153] 5. Terminal: The received result data is decompressed, and the translation result is converted into voice data using a speech synthesis engine (Google Text-to-Speech), and the user is notified that "This sign says 'Business hours: 10:00-22:00'."
[1154] Anomaly detection example
[1155] 1. User: (Unconscious, collapsed)
[1156] 2. Terminal: The camera captures the user's abnormal state in MPEG format and sends the abnormal state data to the server.
[1157] 3. Server: Analyzes abnormal situations using FaceNet and OpenPose and determines whether the user is in danger.
[1158] 4. Server: Contacts emergency services and provides the user's location and video data, and also notifies the user's emergency contacts.
[1159] 5. Terminal: The user is notified that an ambulance has been dispatched using a speech synthesis engine (Google Text-to-Speech).
[1160] Example prompts to input to the generative AI model
[1161] 1. For a translation request: "Please explain how the system works when a user wants the text on a sign translated."
[1162] 2. For anomaly detection: "If the user loses consciousness, explain how the system will detect the anomaly and respond."
[1163] In this way, the system of the present invention can provide users with information quickly and accurately in their daily lives and can also respond to emergencies.
[1164] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1165] Step 1:
[1166] The user speaks to the smart earphones. To request a translation, the user says, "I want the text on the sign translated." Voice data is generated as input.
[1167] Step 2:
[1168] The device uses a microphone to capture sound and generate digital audio data, which is then saved in WAV or MP3 format.
[1169] Step 3:
[1170] The device's first-person camera captures the user's viewpoint and generates video data. Specifically, high-resolution video data in JPEG or MPEG format is generated. The viewpoint video data is generated as input.
[1171] Step 4:
[1172] The audio and video data generated by the device is temporarily stored in local storage and compressed. Specifically, the data size is reduced using GZIP compression. The generated audio and video data is used as input, and a compressed data file is output.
[1173] Step 5:
[1174] The device sends the compressed audio and video data to a server over the Internet. Specifically, the data is sent securely using the HTTPS protocol. The compressed data file is used as input, and a response indicating that it has been sent to the server is output.
[1175] Step 6:
[1176] The server analyzes the received voice data and recognizes the user's intent. Specifically, it converts the voice data into text data using Google Speech-to-Text. The voice data is used as input, and the text data "I would like the words on the sign translated" is output.
[1177] Step 7:
[1178] The server recognizes from the text data that "I want the text on the sign translated" and analyzes the video data. Specifically, it uses OpenCV to extract the text area on the sign from the video data. The video data is used as input, and the coordinate information of the text area is output.
[1179] Step 8:
[1180] The server translates the extracted text into the specified language. Specifically, it uses the Google Translate API to translate the text. The text is used as input, and the translated text "Business hours: 10:00-22:00" is output.
[1181] Step 9:
[1182] The server generates result data containing the translation results and formats it in JSON format. The data is then compressed again using GZIP and sent to the terminal. Specifically, JSON-structured data is generated, compressed, and sent. The translation results are used as input, and a compressed result data file is output.
[1183] Step 10:
[1184] The device unpacks the received result data and uses Google Text-to-Speech to convert the translation results into audio data. Specifically, audio data such as "This sign says 'Business hours: 10:00-22:00'" is generated. The result data is used as input, and audio data is output.
[1185] Step 11:
[1186] The device notifies the user of the generated audio data, specifically by playing the audio through the smart earphones. The audio data is used as input and the user's listening is output.
[1187] (Application example 1)
[1188] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1189] In recent years, the introduction of smart devices and AI technology has progressed in many industries, and efficiency is also required in logistics centers. However, in the current work environment where a large amount of information and real-time instructions are required, employees have to manually check information and receive instructions, which is time-consuming. To solve this problem, a system is needed that allows workers to efficiently receive work instructions using audio and point-of-view video and respond immediately. Therefore, the present invention provides a smart device system that improves work efficiency in logistics centers.
[1190] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1191] In this invention, the server includes means for capturing user voice, means for capturing video from the user's viewpoint, means for transmitting the captured voice data and video data to the server, means for analyzing the voice data by the server to recognize the user's intention, means for analyzing the video data by the server to extract necessary information, means for generating a response based on the analyzed data and transmitting result data to the terminal, means for notifying the user of the transmitted result, and means for combining voice capture and video capture to provide information corresponding to a specific work instruction. This makes it possible for workers at a logistics center to issue voice instructions to provide a pick list, and for camera video to be analyzed to notify inventory information in real time.
[1192] Key Word Definitions
[1193] An "audio capture means" is a device or method that receives user-uttered audio and converts it into digital data.
[1194] A "means for capturing viewpoint images" is a device or method for acquiring images seen from the user's field of view in real time and recording them as digital data.
[1195] The "means for transmitting audio data and video data to a server" refers to a device or method for transferring the captured audio data and video data to a server via a communication network such as the Internet.
[1196] "Means for analyzing voice data and recognizing user intent" refers to algorithms or software that analyze received voice data and understand what the user is requesting.
[1197] "Means for analyzing video data and extracting necessary information" refers to algorithms or software that analyze received video data and identify and extract specific information (for example, text information or object information).
[1198] The "means for generating a response and transmitting result data to the terminal" refers to a device or method that generates an appropriate response based on the analyzed data and forwards the result to the user terminal.
[1199] "Means for notifying the user of the transmitted results" refers to a device or method for notifying the user of the transmitted results data at the terminal, including, for example, audio notification or visual display.
[1200] "Means for providing information corresponding to specific work instructions" refers to a method or device that combines audio and video capture to provide a user with specific instructions or information for a specific task or work.
[1201] A "pick list" is a list of instructions provided to workers for picking and sorting specific items at a distribution center.
[1202] "Inventory information" refers to data regarding the type, quantity, location, etc. of products stored in warehouses and logistics centers.
[1203] MODE FOR CARRYING OUT THE INVENTION
[1204] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system is particularly suitable for improving work efficiency in logistics centers. Specific embodiments for implementing this system are described below.
[1205] System configuration
[1206] This system consists of the following main components:
[1207] 1. User device (smart earphones and first-person camera)
[1208] 2. Server (AI processing unit)
[1209] 3. Internet connection
[1210] User terminal
[1211] The user terminal includes a microphone for capturing the user's voice and a first-person camera for capturing the user's point of view. The terminal also includes a communication module for transmitting data to a server.
[1212] server
[1213] The server includes an AI processing unit for analyzing the received audio and video data. This allows the server to recognize the user's intent and execute processing to provide the necessary information. The server also returns the resulting data to the device, which notifies the user of the results.
[1214] System Operation
[1215] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[1216] 1. Audio capture and transmission
[1217] The user speaks to the smart earphones, saying, "Tell me the next picklist." The device's microphone captures this voice and generates audio data. At the same time, a camera capturing the user's point of view captures the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed, and then sent to a server via the Internet.
[1218] 2. Analysis and processing on the server
[1219] The server analyzes the received voice data and recognizes the user's intent. For example, it extracts the intent "next pick list" from the voice. At the same time, it analyzes the video data and identifies shelf information, for example.
[1220] 3. Generating and returning result data
[1221] The server generates a picklist based on the analysis results, compresses this data again, and sends it to the terminal.
[1222] 4. Notice to Users
[1223] The terminal decompresses the received result data and notifies the user by voice, for example, by giving instructions such as "Please take product B from shelf A3."
[1224] Specific examples
[1225] Picklist Request Example
[1226] 1. User: "What's the next picklist?"
[1227] 2. Device: The microphone captures audio and the camera captures video of the surroundings. This data is sent to the server.
[1228] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the shelf information, then generates a pick list.
[1229] 4. Server: Sends the picklist to the terminal.
[1230] 5. Terminal: The picklist is announced to the user via voice: "Please pick up product B from shelf A3."
[1231] In this way, the system of the present invention can provide pick lists by voice instructions from logistics center workers and can analyze camera images to notify inventory information in real time. This system is expected to significantly improve work efficiency at logistics centers.
[1232] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1233] System program processing flow
[1234] Step 1:
[1235] The user speaks to the smart earphones and gives a voice command such as "Tell me the next picklist." This voice is captured by the microphone and generated as voice data.
[1236] Input: User's voice
[1237] Output: Audio data
[1238] How it works: The microphone in the smart earphones captures the user's voice and converts it into digital data.
[1239] Step 2:
[1240] The device's first-person camera captures images from the user's perspective and generates video data.
[1241] Input: Video from the user's point of view
[1242] Output: Video data
[1243] Specific operation: The device's camera captures video from the user's perspective in real time and stores it as digital data.
[1244] Step 3:
[1245] The captured audio and video data is temporarily stored in the terminal, compressed, and then transmitted to a server via the Internet.
[1246] Input: Audio data, video data
[1247] Output: Data sent to the server
[1248] Specific operation: Audio and video data is stored in the device's local storage, the data is compressed using a compression algorithm, and then sent to a server via an Internet connection.
[1249] Step 4:
[1250] The server analyzes the received voice data and recognizes the user's intent, for example, extracting the intent "next picklist."
[1251] Input: Audio data
[1252] Output: User intent (text format)
[1253] Specific operation: The server's AI processing unit uses a speech recognition engine (e.g., Google's speech recognition API) to convert the voice data into text and analyze the user's intent.
[1254] Step 5:
[1255] At the same time, the server analyzes the video data to identify, for example, shelf information.
[1256] Input: Video data
[1257] Output: Analysis results (shelf information)
[1258] Specific operation: Use the server's video analysis engine (such as OpenCV) to extract and identify specific information (e.g., shelf labels) from the video data.
[1259] Step 6:
[1260] The server generates a picklist based on the analysis results, compresses this data again, and sends it to the terminal.
[1261] Input: User intent, shelf information
[1262] Output: Picklist data
[1263] Specific operation: The generated picklist is compressed for transmission to the terminal, and then transmitted to the terminal via an Internet connection.
[1264] Step 7:
[1265] The terminal decompresses the received result data and notifies the user by voice, for example, by giving instructions such as "Please take product B from shelf A3."
[1266] Input: Compressed result data
[1267] Output: Audio notification
[1268] Specific operation: The device decompresses the compressed data and notifies the user by voice using a speech synthesis engine (e.g., Google TTS).
[1269] This series of processing steps allows workers at the logistics center to efficiently receive work instructions and respond immediately.
[1270] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1271] This invention is a smart earphone system that provides users with necessary information in real time in their daily lives and recognizes their emotions to provide appropriate feedback. This system captures audio and video from various viewpoints and analyzes the data on a server to generate a variety of responses. Furthermore, by combining it with an emotion engine, it is possible to provide adaptive services according to the user's emotions.
[1272] System configuration
[1273] This system consists of the following main components:
[1274] 1. User device (smart earphones and first-person camera)
[1275] 2. Server (AI processing unit and emotion engine)
[1276] 3. Internet connection
[1277] User terminal
[1278] The user terminal includes a microphone to capture the user's voice, a first-person camera to capture the user's point of view, and a communication module to transmit this data to a server via the Internet.
[1279] server
[1280] The server includes an AI processing unit for analyzing the received audio and video data. It is connected to an emotion engine, which is capable of recognizing the user's emotions from the received data. The server then generates appropriate information and responses based on the user's intentions and emotions, and notifies the user via the device.
[1281] System Operation
[1282] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[1283] Audio capture and transmission
[1284] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and converts it into digital audio data. At the same time, a first-person camera captures video from the user's point of view and converts it into digital video data. This data is temporarily stored in local storage, then compressed and sent to a server using a secure protocol.
[1285] Server parsing and processing
[1286] The server analyzes the received voice data and converts it into text data using a speech recognition engine. For example, it recognizes the intent of "I want it translated." At the same time, it analyzes the video data, identifies the text area on the sign, and extracts the text information. It then uses an emotion engine to recognize the user's emotions from the voice and video data, generating information to indicate the urgency of the request and provide adaptive feedback. The server then translates the text information into the specified language using a translation engine.
[1287] Generate and return result data
[1288] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[1289] User Notification
[1290] The device receives and decompresses the resulting data. It then notifies the user of the translation results and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[1291] Specific examples
[1292] Translation request example
[1293] 1. User: "I want the text on the sign translated."
[1294] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[1295] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the characters on the sign. The emotion engine analyzes the user's urgency and emotion from the voice and video. The translation engine then translates the characters into the specified language.
[1296] 4. Server: Sends translation results and emotion recognition results to the device.
[1297] 5. Terminal: The user is notified of the translation result and emotional feedback via voice, such as "This sign says 'Business hours: 10:00-22:00'. Has your problem been resolved?"
[1298] Anomaly detection example
[1299] 1. User: (Unconscious, collapsed)
[1300] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[1301] 3. Server: Analyzes the video data and anomaly detection data and determines whether the user is in danger.
[1302] 4. Server: Generates and sends data to contact emergency services and provide them with the necessary information, as well as to notify registered emergency contacts.
[1303] 5. Terminal: The user is notified by voice that the emergency call has been completed. The user is told, "An ambulance has been dispatched. Please wait a moment."
[1304] In this way, the system of the present invention can provide information quickly and accurately in the user's daily life and provide adaptive feedback according to the user's emotions.
[1305] The processing flow will be explained below.
[1306] Specific processing steps
[1307] Example 1: Processing a translation request
[1308] Step 1:
[1309] The user speaks to the smart earphones and says, "I want the text on the sign translated."
[1310] Step 2:
[1311] The device's microphone captures the user's voice, which is then converted into a digital format.
[1312] Step 3:
[1313] The device's first-person camera captures video from the user's point of view, and the video data is converted into a digital format.
[1314] Step 4:
[1315] The audio and video data captured by the device is temporarily stored in local storage.
[1316] Step 5:
[1317] The terminal transmits the audio and video data to the server using a secure protocol.
[1318] Step 6:
[1319] The server receives the voice data and converts it into text data using a voice recognition engine. For example, a command such as "Please translate the text on the sign" is converted into text information.
[1320] Step 7:
[1321] The server receives the video data and uses computer vision technology to identify text regions in the video, then extracts the identified text information.
[1322] Step 8:
[1323] The server combines the results of speech recognition and video analysis to understand the user's intent. In this case, it recognizes the intent to "translate the text on the sign."
[1324] Step 9:
[1325] The emotion engine analyzes audio and video data to recognize the user's emotions. For example, it determines that the user is feeling anxious.
[1326] Step 10:
[1327] The server translates the text information into the specified language using a translation engine, for example, translating a Japanese sign into English.
[1328] Step 11:
[1329] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[1330] Step 12:
[1331] The terminal receives and decompresses the resulting data.
[1332] Step 13:
[1333] The device will then provide the user with a voice message with the result data, for example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[1334] Example 2: Emergency call due to abnormality detection
[1335] Step 1:
[1336] The user falls into an abnormal state such as losing consciousness.
[1337] Step 2:
[1338] The device's first-person camera captures the user's unusual situation.
[1339] Step 3:
[1340] The device acquires anomaly detection data from the biometric sensor and temporarily stores it in local storage.
[1341] Step 4:
[1342] The video data and anomaly detection data captured by the device are sent to the server using a secure protocol.
[1343] Step 5:
[1344] The server analyzes the received video data and anomaly detection data.
[1345] Step 6:
[1346] The server determines from the analysis results that the user is in danger.
[1347] Step 7:
[1348] The emotion engine analyzes the data and determines that the user is in an emergency.
[1349] Step 8:
[1350] The server automatically contacts the nearest emergency services and provides them with the necessary information, such as the user's current location and condition.
[1351] Step 9:
[1352] The server generates data to notify the registered emergency contacts and transmits this to the terminal.
[1353] Step 10:
[1354] The device receives the emergency call results and provides the information to relatives and friends via voice or other notification means.
[1355] Step 11:
[1356] The terminal notifies the user by voice that the ambulance has been dispatched, for example, by saying, "The ambulance has been dispatched. Please wait a moment."
[1357] The above are the specific processing steps in the system of the present invention.
[1358] Example 2
[1359] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1360] Conventional information provision systems have had difficulty efficiently analyzing the user's voice and visual perspective and providing appropriate information and feedback in real time. They also lacked the technology to recognize the user's emotions and generate appropriate responses accordingly. Furthermore, there was no mechanism for detecting abnormal situations and responding quickly, which created problems in ensuring the user's safety.
[1361] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing voice data and converting it into text data using a voice recognition engine, means for analyzing video data and extracting information using a video analysis engine, and means for recognizing the user's emotions from the voice data and video data using an emotion engine. This makes it possible to analyze the user's voice and video from their viewpoint in real time and provide adaptive information and feedback according to the user's intentions and emotions. Furthermore, in an emergency, abnormalities can be detected and emergency services can be quickly contacted, thereby ensuring the user's safety.
[1362] "User" refers to an individual who uses the system and receives information.
[1363] "Means for capturing audio" refers to a microphone and its peripheral devices for converting a user's speech into digital data.
[1364] The term "means for capturing an image of a viewpoint" refers to a camera and its peripheral devices for converting an image seen from the user's viewpoint into digital data.
[1365] "Means for transmitting captured audio and video data to a server" refers to a communications module for transmitting digital data to a server via the Internet or other network.
[1366] "Means for analyzing voice data and converting it into text data using a voice recognition engine" refers to voice recognition technology for processing voice data and converting it into text information.
[1367] "Means for analyzing video data and extracting information using a video analytics engine" refers to video analytics technology for processing video data to extract specific information.
[1368] "Means for recognizing a user's emotions from audio data and video data using an emotion engine" refers to a technology for analyzing a user's emotions by understanding the characteristics of audio and video.
[1369] "Means for generating a response based on the analyzed data and transmitting the resulting data to the terminal" refers to algorithms and communication modules for generating an appropriate response based on the analyzed information and transmitting the response to the terminal.
[1370] "Means for notifying the user of the transmitted results" refers to a speaker or display on the terminal for displaying or notifying the user of the results by voice.
[1371] "Text data" refers to the text information of a user's speech converted by a voice recognition engine.
[1372] "Means for extracting information" refers to the processes and algorithms used by the video analytics engine to extract specific information.
[1373] "Means for recognizing emotions" refers to a process for analyzing and identifying a user's emotions using an emotion engine.
[1374] "Results data" refers to data that includes response information generated by a server.
[1375] This invention relates to a system that provides users with necessary information in real time in their daily lives and recognizes their emotions to provide appropriate feedback. This system captures audio and video from various viewpoints and analyzes the data on a computer server, enabling it to generate a variety of responses. Furthermore, by combining it with an emotion engine, it becomes possible to provide adaptive services according to the user's emotions.
[1376] System configuration
[1377] This system consists of the following main components:
[1378] 1. User device (smart earphones and first-person camera)
[1379] 2. Server (AI processing unit and emotion engine)
[1380] 3. Internet connection
[1381] User terminal
[1382] The user terminal includes a microphone to capture the user's voice, a first-person camera to capture the user's point of view, and a communication module to transmit this data to a server via the Internet.
[1383] server
[1384] The server includes an AI processing unit for analyzing the received audio and video data. It is connected to an emotion engine, which is capable of recognizing the user's emotions from the received data. The server then generates appropriate information and responses based on the user's intentions and emotions, and notifies the user via the device.
[1385] Specific operation example
[1386] Audio capture and transmission
[1387] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and converts it into digital audio data. At the same time, a first-person camera captures video from the user's point of view and converts it into digital video data. This data is temporarily stored in local storage, then compressed and sent to a server using a secure protocol.
[1388] Server parsing and processing
[1389] The server analyzes the received voice data and converts it into text data using a speech recognition engine. For example, it recognizes the intent of "I want it translated." At the same time, it analyzes the video data, identifies the text area on the sign, and extracts the text information. It then uses an emotion engine to recognize the user's emotions from the voice and video data, generating information to indicate the urgency of the request and provide adaptive feedback. The server then translates the text information into the specified language using a translation engine.
[1390] Generate and return result data
[1391] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[1392] User Notification
[1393] The device receives and decompresses the resulting data. It then notifies the user of the translation results and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[1394] Specific examples
[1395] Translation request example
[1396] 1. User: "I want the text on the sign translated."
[1397] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[1398] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the characters on the sign. The emotion engine analyzes the user's urgency and emotion from the voice and video. The translation engine then translates the characters into the specified language.
[1399] 4. Server: Sends translation results and emotion recognition results to the device.
[1400] 5. Terminal: The user is notified of the translation result and emotional feedback via voice, such as "This sign says 'Business hours: 10:00-22:00'. Has your problem been resolved?"
[1401] Anomaly detection example
[1402] 1. User: (Unconscious, collapsed)
[1403] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[1404] 3. Server: Analyzes the video data and anomaly detection data and determines whether the user is in danger.
[1405] 4. Server: Generates and sends data to contact emergency services and provide them with the necessary information, as well as to notify registered emergency contacts.
[1406] 5. Terminal: The user is notified by voice that the emergency call has been completed. The user is told, "An ambulance has been dispatched. Please wait a moment."
[1407] In this way, the system of the present invention can provide information quickly and accurately in the user's daily life and provide adaptive feedback according to the user's emotions.
[1408] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1409] Step 1:
[1410] The user speaks to the smart earphones, saying, "I want the text on the sign to be translated." The user's voice becomes the input.
[1411] Step 2:
[1412] The device's microphone captures the user's voice and converts it from an analog audio signal into digital audio data, which is then output.
[1413] Step 3:
[1414] The device's first-person camera captures video from the user's point of view and converts the analog video signal into digital video data, which is then output.
[1415] Step 4:
[1416] The device temporarily stores digital audio and video data in local storage, which acts as a temporary buffer.
[1417] Step 5:
[1418] The device compresses the stored data (e.g., in gzip format) to make it easier to analyze, and the compressed data is output.
[1419] Step 6:
[1420] The device sends the compressed data to the server using a secure protocol (e.g., HTTPS), which then becomes the server's input.
[1421] Step 7:
[1422] The server analyzes the received voice data using a voice recognition engine (e.g., general voice recognition technology) and converts the digital voice data into text data, which is then output.
[1423] Step 8:
[1424] The server analyzes the text data and recognizes the user's intent (e.g., a request to "translate"). The recognition result becomes the input for the next step.
[1425] Step 9:
[1426] The server analyzes the received video data using a video analysis engine (e.g., general video recognition technology) to identify the text area on the sign. The identified text area becomes the output.
[1427] Step 10:
[1428] The server extracts text information from the specified text area, and this text information becomes the output.
[1429] Step 11:
[1430] The server uses an emotion engine (e.g., general emotion analysis technology) to recognize the user's emotion from the audio and video data. The emotion recognition result is the output.
[1431] Step 12:
[1432] The server generates a response based on the analysis results (user intent, text information, and emotion recognition results). The generated response becomes the output.
[1433] Step 13:
[1434] The server compresses the response data (e.g., in gzip format) and sends it to the terminal. The sent data becomes the terminal's input.
[1435] Step 14:
[1436] The device decompresses the received data and provides the translation results and feedback to the user in a voice message with a tone and content that adapts to the user's emotions, such as, "This sign says 'Business hours: 10:00-22:00'. Is there anything I can help you with?"
[1437] (Application example 2)
[1438] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1439] Conventional work support systems in factories often provide limited information and basic feedback, lacking adaptive support based on the worker's emotions and situation. They also lack early detection and rapid response to abnormal situations during work. This makes it difficult to improve work efficiency and ensure safety.
[1440] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1441] In this invention, the server includes a means for recognizing the user's emotions and generating an appropriate response based on the emotions, a means for transmitting work instructions and safety guidelines in real time within the factory, and a means for detecting abnormalities from received video data and automatically contacting emergency services. This enables adaptive support based on the emotions of workers, as well as early detection and rapid response to abnormal situations during work.
[1442] "User" refers to an individual who uses the system, particularly a worker who performs work within a factory.
[1443] "Audio data" refers to digital data that captures audio information uttered by a user.
[1444] "Visual data" refers to digital data that contains video information captured from a user's point of view.
[1445] "Server" refers to a computing device that analyzes the transmitted audio and visual data, extracts the necessary information, and generates a response for the user.
[1446] "Emotion recognition" refers to the process of analyzing a user's audio and visual data to identify the user's emotional state.
[1447] "Work instructions" refers to instruction information that conveys specific work content and procedures to users within a factory.
[1448] "Safety guidelines" refers to instruction documents that clearly state the safety measures and precautions that should be observed during work.
[1449] "Anomaly detection" refers to the process of analyzing received video data to identify situations that deviate from normal working conditions or emergencies.
[1450] "Emergency services" refers to external support services such as ambulances, fire departments, and police departments that respond to emergency situations during work.
[1451] This invention is a smart earphone system that provides real-time support to workers (users) as a work support system in factories. The system captures and analyzes the user's voice and visual data to recognize emotions and provide adaptive work instructions and safety guidelines. It also has the ability to detect abnormalities and automatically contact emergency services.
[1452] System configuration
[1453] This system consists of the following main components:
[1454] 1. User Device
[1455] The user terminal includes a smart earphone and a first-person camera worn by the worker. The terminal is equipped with a microphone for capturing the user's voice, a camera for capturing first-person perspective video, and a communication module for transmitting the data to a server.
[1456] 2. Server
[1457] The server is equipped with an AI processing unit and emotion engine for receiving and analyzing voice and visual data. It has the ability to analyze voice data to recognize the user's intentions and visual data to extract necessary information. The emotion engine also recognizes the user's emotions and generates adaptive responses according to those emotions.
[1458] System Operation
[1459] When a user speaks to the smart earbuds, the system works as follows:
[1460] 1. Audio and visual data capture
[1461] When a user requests a work instruction, the smart earphones' microphone captures audio and the camera captures video from the user's point of view, and these digital data are temporarily stored in local storage before being transmitted to a server using a secure protocol.
[1462] 2. Analysis and processing on the server
[1463] The server analyzes the received voice data and converts it into text data using a speech recognition engine. At the same time, it analyzes the visual data to extract the necessary information. For example, if a user requests, "Please tell me the next step in the process," the server recognizes the user's intention and generates information about the next step.
[1464] 3. Emotion Recognition and Response Generation
[1465] The emotion engine recognizes the user's emotions from audio and visual data, generating information to determine the urgency of the request and provide adaptive feedback. For example, if the user is feeling stressed, it will provide a safety warning.
[1466] 4. Generating and returning result data
[1467] The server generates result data including analysis results, emotion recognition results, and work instructions, compresses it, and sends it to the terminal.
[1468] 5. Notice to Users
[1469] The device receives and decompresses the resulting data. It then provides the user with work instructions and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "Next, install part A. Safety first."
[1470] Specific examples
[1471] Work instruction example
[1472] The user speaks to the system, saying, "Please tell me the next steps," and the first-person camera captures the work environment. The system analyzes the audio and video, generates appropriate work instructions, and notifies the user by voice.
[1473] Example prompts to input to a generative AI model:
[1474] The user speaks to the smart earphones, "What are the next steps?", and the first-person camera captures the work environment. What steps can be taken to analyze the audio and video, generate appropriate work instructions, and provide voice prompts?
[1475] In this way, the system of the present invention can provide information quickly and accurately to workers in their daily work, and can provide adaptive feedback according to the user's emotions.
[1476] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1477] Step 1:
[1478] This is the step where the user speaks instructions to the smart earphone. For example, the user might say, "Please tell me the next step." This voice input triggers the start of the next processing step.
[1479] Step 2:
[1480] This is the step where the device captures the audio and converts it into digital audio data.
[1481] Input: User's voice
[1482] Output: Digital audio data
[1483] The device's microphone captures the user's voice, and an internal processor converts this voice into digital audio data.
[1484] Step 3:
[1485] The terminal simultaneously captures the visual data and converts it into digital video data.
[1486] Input: Video from the user's point of view
[1487] Output: Digital video data
[1488] A first-person camera mounted on the device captures video from the user's point of view and converts this video into digital data.
[1489] Step 4:
[1490] This is a step in which the terminal transmits the captured audio and video data to the server.
[1491] Input: Digital audio data, digital video data
[1492] Output: Send data to the server
[1493] A communication module in the terminal transmits the captured audio and video data to a server over a secure protocol.
[1494] Step 5:
[1495] This is the step where the server analyzes the received voice data and recognizes the user's intention.
[1496] Input: Digital audio data
[1497] Output: Text data and intent recognition results
[1498] The server's speech recognition engine converts the voice data into text data, analyzes it, and identifies the user's intent.
[1499] Step 6:
[1500] This is the step where the server analyzes the received visual data and extracts the necessary information.
[1501] Input: Digital video data
[1502] Output: Extracted information
[1503] The visual analytics engine analyzes the video data and identifies necessary information (such as the state of the work environment or specific objects).
[1504] Step 7:
[1505] This is the step where the server recognizes the user's emotions from the audio data and visual data.
[1506] Input: Digital audio data, digital video data
[1507] Output: User's emotional state
[1508] The emotion engine analyzes audio and video to identify the user's emotional state.
[1509] Step 8:
[1510] This is the step in which the server generates a response based on the analyzed data, creates result data, and sends it to the terminal.
[1511] Input: Intention recognition results, extracted information, emotional state
[1512] Output: Result data
[1513] The server generates an appropriate response based on the user's intention, information about the working environment, and emotional state, and creates result data, which is then compressed and sent to the terminal.
[1514] Step 9:
[1515] This is the step in which the terminal decompresses the received result data and notifies the user.
[1516] Input: Compressed result data
[1517] Output: Audio notification
[1518] The terminal decompresses the received result data and notifies the user by voice of appropriate work instructions and precautions, such as "Next, please install part A. Safety first."
[1519] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1520] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1521] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1522] [Fourth embodiment]
[1523] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1524] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1525] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1526] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1527] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1528] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1529] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1530] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1531] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1532] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1533] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1534] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1535] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1536] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system has a series of functions: capturing audio, capturing viewpoint video, analyzing data, generating results, and notifying the user of those results. Specific embodiments of the system are described below.
[1537] System configuration
[1538] This system consists of the following main components:
[1539] 1. User device (smart earphones and first-person camera)
[1540] 2. Server (AI processing unit)
[1541] 3. Internet connection
[1542] User terminal
[1543] The user terminal includes a microphone for capturing the user's voice and a first-person camera for capturing the user's point of view. The terminal also includes a communication module for transmitting data to a server.
[1544] server
[1545] The server includes an AI processing unit for analyzing the received audio and video data. This allows the server to recognize the user's intent and execute processing to provide the necessary information. The server also returns the resulting data to the device, which notifies the user of the results.
[1546] System Operation
[1547] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[1548] Audio capture and transmission
[1549] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and generates audio data. At the same time, a camera capturing the user's point of view captures video of the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed, and then sent to a server via the Internet.
[1550] Server parsing and processing
[1551] The server analyzes the received voice data and recognizes the user's intent. For example, it extracts the intent "signboard" and "please translate" from the voice. At the same time, it analyzes the video data and identifies the text area on the signboard. The identified text information is translated into the specified language by the translation engine.
[1552] Generate and return result data
[1553] The server generates result data including the translation result, compresses this data again, and transmits it to the terminal.
[1554] User Notification
[1555] The device decompresses the received result data and notifies the user by voice, for example, "This sign says 'Business hours: 10:00-22:00'."
[1556] Specific examples
[1557] Translation request example
[1558] 1. User: "I want the text on the sign translated."
[1559] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[1560] 3. Server: The speech recognition engine converts the voice data into text, the video analysis engine recognizes the characters on the sign, and the translation engine then translates the characters into the specified language.
[1561] 4. Server: Sends the translation results to the device.
[1562] 5. Terminal: The translation result is notified to the user by voice.
[1563] Anomaly detection example
[1564] 1. User: (Unconscious, collapsed)
[1565] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[1566] 3. Server: Analyzes abnormal situations and determines that the user is in danger.
[1567] 4. Server: Contacts emergency services and provides them with the necessary information, and also notifies emergency contacts.
[1568] 5. Terminal: The user is notified by voice that an ambulance has been dispatched.
[1569] In this way, the system of the present invention can provide users with information quickly and accurately in their daily lives and can also respond to emergencies.
[1570] The processing flow will be explained below.
[1571] Specific processing steps of the program
[1572] Example 1: Processing a translation request
[1573] Step 1:
[1574] The user speaks to the smart earphones and says, "I want the text on the sign translated."
[1575] Step 2:
[1576] The device's microphone captures the user's voice, which is then converted into digital data.
[1577] Step 3:
[1578] The device's camera captures an image of the sign from the user's point of view, and the image data is also converted into digital data.
[1579] Step 4:
[1580] The audio and video data captured by the device is temporarily stored in local storage.
[1581] Step 5:
[1582] The terminal transmits audio and video data to the server using a secure protocol.
[1583] Step 6:
[1584] The server receives the voice data and uses a speech recognition engine to convert the digital speech into text, for example, recognizing a command such as "Please translate the text on the sign."
[1585] Step 7:
[1586] The server receives the video data and uses computer vision techniques to identify text areas in the video, for example, to extract text on a sign.
[1587] Step 8:
[1588] The server combines the results of voice recognition and video analysis to understand the user's intent (in this case, "translating the text on the sign").
[1589] Step 9:
[1590] The server uses a translation engine to translate the sign text into the specified language, for example, translating Japanese text into English.
[1591] Step 10:
[1592] The server generates result data including the translation result, compresses it, and transmits it to the terminal.
[1593] Step 11:
[1594] The terminal receives and decompresses the resulting data.
[1595] Step 12:
[1596] The device will then notify the user of the translation results by voice, for example, saying, "This sign says 'Business hours: 10:00-22:00'."
[1597] Example 2: Emergency call due to abnormality detection
[1598] Step 1:
[1599] The user falls into an abnormal state such as losing consciousness.
[1600] Step 2:
[1601] The device's camera captures the user's unusual postures and situations.
[1602] Step 3:
[1603] The device acquires anomaly detection data from biometric sensors and other devices and temporarily stores it in local storage.
[1604] Step 4:
[1605] The video data and anomaly detection data captured by the device are sent to the server using a secure protocol.
[1606] Step 5:
[1607] The server analyzes the received video data and anomaly detection data.
[1608] Step 6:
[1609] The server determines from the analysis results that the user is in danger.
[1610] Step 7:
[1611] The server automatically contacts the nearest emergency services and provides them with the necessary information, such as the user's current location and condition.
[1612] Step 8:
[1613] The server generates data to notify the registered emergency contacts and transmits this to the terminal.
[1614] Step 9:
[1615] The device receives the emergency call results and provides the information to relatives and friends via voice or other notification means.
[1616] Step 10:
[1617] The terminal notifies the user by voice that the ambulance has been dispatched, for example, by saying, "The ambulance has been dispatched. Please wait a moment."
[1618] The above are the specific processing steps in the system of the present invention.
[1619] Example 1
[1620] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1621] Current speech recognition and translation systems can sometimes make it difficult for users to quickly and accurately obtain the information they need. Furthermore, there is a lack of systems that can monitor users' status in real time and provide appropriate responses in emergencies. Therefore, there is a need for systems that can provide users with quick and accurate information in everyday life and respond appropriately in emergencies.
[1622] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1623] In this invention, the server includes means for analyzing voice data using a voice recognition engine to recognize the user's intention, means for analyzing video data using image analysis means to extract necessary information, and means for translating the extracted information into a specified language using a translation engine, thereby making it possible to recognize the user's intention from the voice data, extract necessary information from the video data, and further translate it.
[1624] An "audio capturing means" is a device or software for collecting and digitizing a user's voice.
[1625] The "means for capturing viewpoint images" is a device or software for collecting images based on the user's viewpoint.
[1626] A "compression means" is a device or software that compresses audio and video data to reduce its size.
[1627] The "means for transmitting to the server" is a device or software for transmitting compressed audio data and video data to the server via the Internet.
[1628] "Means for analyzing and recognizing the user's intent using a voice recognition engine" refers to software that analyzes voice data on a server, converts it into text data, and understands the user's requests and intent.
[1629] "Means for analyzing and extracting necessary information using image analysis means" refers to software that analyzes video data on a server and identifies and extracts important information from it.
[1630] The "means for translating into a specified language using a translation engine" is software for translating extracted information into another language.
[1631] The "means for generating a response and transmitting the resulting data to the terminal" refers to software and communication devices for generating a response based on the analyzed and translated data and transmitting the resulting data to the terminal.
[1632] The "means for notifying the user using a voice synthesis engine" refers to software and a playback device for converting the result data into voice data and notifying the user.
[1633] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system has a series of functions: capturing audio, capturing viewpoint video, analyzing data, generating results, and notifying the user of those results. Specific embodiments of the system are described below.
[1634] System configuration
[1635] This system consists of the following main components:
[1636] 1. User device (smart earphones and first-person camera)
[1637] 2. Server (AI processing unit)
[1638] 3. Internet connection
[1639] User terminal
[1640] The user terminal includes a microphone to capture the user's voice and a first-person camera to capture the user's point of view. The terminal also has a communications module for transmitting data to a server. A digital microphone is used to capture the voice, and a high-resolution camera is used to capture the point of view. Data compression is performed using MP3 format for audio data and JPEG or MPEG format for video data.
[1641] server
[1642] The server includes an AI processing unit for analyzing the received audio and video data. For example, Google Speech-to-Text is used as the speech recognition engine, and OpenCV is used for image analysis. Furthermore, the Google Translate API is used as the translation engine. The analyzed and translated data is stored in JSON format and sent back to the device after GZIP compression. The device then uses a speech synthesis engine (e.g., Google Text-to-Speech) to notify the user.
[1643] System Operation
[1644] In an embodiment of the system, when a user speaks to the smart earphones, the system operates as follows: When the user speaks to the smart earphones, saying, "I want the words on the sign translated," the digital microphone on the user device captures this voice and generates audio data. At the same time, the first-person camera captures video of the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed into MP3 and JPEG formats, and then transmitted to a server via the Internet.
[1645] The server analyzes the received voice data using Google Speech-to-Text to recognize the user's intent. For example, it extracts the intents "signboard" and "translate" from the voice. At the same time, it analyzes the video data using OpenCV to identify the text area on the signboard. The identified text information is translated into the specified language using the Google Translate API.
[1646] The server generates result data containing the translation results, composes this data in JSON format, compresses it with GZIP, and sends it to the device. The device then decompresses the received result data and uses Google Text-to-Speech to notify the user aloud. For example, "This sign says 'Business hours: 10:00-22:00'."
[1647] Specific examples
[1648] Translation request example
[1649] 1. User: "I want the text on the sign translated."
[1650] 2. Terminal: The microphone captures the audio in MP3 format, and the camera takes a video of the sign in JPEG format. These data are sent to the server.
[1651] 3. Server: Uses a speech recognition engine (Google Speech-to-Text) to convert the voice data into text data such as "Please translate the text on the sign." OpenCV is used to extract the text area from the video of the sign, and the Google Translate API is used to translate it into the specified language.
[1652] 4. Server: Generates result data in JSON format, including the translation result "Business hours: 10:00-22:00", compresses it with GZIP, and sends it to the terminal.
[1653] 5. Terminal: The received result data is decompressed, and the translation result is converted into voice data using a speech synthesis engine (Google Text-to-Speech), and the user is notified that "This sign says 'Business hours: 10:00-22:00'."
[1654] Anomaly detection example
[1655] 1. User: (Unconscious, collapsed)
[1656] 2. Terminal: The camera captures the user's abnormal state in MPEG format and sends the abnormal state data to the server.
[1657] 3. Server: Analyzes abnormal situations using FaceNet and OpenPose and determines whether the user is in danger.
[1658] 4. Server: Contacts emergency services and provides the user's location and video data, and also notifies the user's emergency contacts.
[1659] 5. Terminal: The user is notified that an ambulance has been dispatched using a speech synthesis engine (Google Text-to-Speech).
[1660] Example prompts to input to the generative AI model
[1661] 1. For a translation request: "Please explain how the system works when a user wants the text on a sign translated."
[1662] 2. For anomaly detection: "If the user loses consciousness, explain how the system will detect the anomaly and respond."
[1663] In this way, the system of the present invention can provide users with information quickly and accurately in their daily lives and can also respond to emergencies.
[1664] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1665] Step 1:
[1666] The user speaks to the smart earphones. To request a translation, the user says, "I want the text on the sign translated." Voice data is generated as input.
[1667] Step 2:
[1668] The device uses a microphone to capture sound and generate digital audio data, which is then saved in WAV or MP3 format.
[1669] Step 3:
[1670] The device's first-person camera captures the user's viewpoint and generates video data. Specifically, high-resolution video data in JPEG or MPEG format is generated. The viewpoint video data is generated as input.
[1671] Step 4:
[1672] The audio and video data generated by the device is temporarily stored in local storage and compressed. Specifically, the data size is reduced using GZIP compression. The generated audio and video data is used as input, and a compressed data file is output.
[1673] Step 5:
[1674] The device sends the compressed audio and video data to a server over the Internet. Specifically, the data is sent securely using the HTTPS protocol. The compressed data file is used as input, and a response indicating that it has been sent to the server is output.
[1675] Step 6:
[1676] The server analyzes the received voice data and recognizes the user's intent. Specifically, it converts the voice data into text data using Google Speech-to-Text. The voice data is used as input, and the text data "I would like the words on the sign translated" is output.
[1677] Step 7:
[1678] The server recognizes from the text data that "I want the text on the sign translated" and analyzes the video data. Specifically, it uses OpenCV to extract the text area on the sign from the video data. The video data is used as input, and the coordinate information of the text area is output.
[1679] Step 8:
[1680] The server translates the extracted text into the specified language. Specifically, it uses the Google Translate API to translate the text. The text is used as input, and the translated text "Business hours: 10:00-22:00" is output.
[1681] Step 9:
[1682] The server generates result data containing the translation results and formats it in JSON format. The data is then compressed again using GZIP and sent to the terminal. Specifically, JSON-structured data is generated, compressed, and sent. The translation results are used as input, and a compressed result data file is output.
[1683] Step 10:
[1684] The device unpacks the received result data and uses Google Text-to-Speech to convert the translation results into audio data. Specifically, audio data such as "This sign says 'Business hours: 10:00-22:00'" is generated. The result data is used as input, and audio data is output.
[1685] Step 11:
[1686] The device notifies the user of the generated audio data, specifically by playing the audio through the smart earphones. The audio data is used as input and the user's listening is output.
[1687] (Application example 1)
[1688] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1689] In recent years, the introduction of smart devices and AI technology has progressed in many industries, and efficiency is also required in logistics centers. However, in the current work environment where a large amount of information and real-time instructions are required, employees have to manually check information and receive instructions, which is time-consuming. To solve this problem, a system is needed that allows workers to efficiently receive work instructions using audio and point-of-view video and respond immediately. Therefore, the present invention provides a smart device system that improves work efficiency in logistics centers.
[1690] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1691] In this invention, the server includes means for capturing user voice, means for capturing video from the user's viewpoint, means for transmitting the captured voice data and video data to the server, means for analyzing the voice data by the server to recognize the user's intention, means for analyzing the video data by the server to extract necessary information, means for generating a response based on the analyzed data and transmitting result data to the terminal, means for notifying the user of the transmitted result, and means for combining voice capture and video capture to provide information corresponding to a specific work instruction. This makes it possible for workers at a logistics center to issue voice instructions to provide a pick list, and for camera video to be analyzed to notify inventory information in real time.
[1692] Key Word Definitions
[1693] An "audio capture means" is a device or method that receives user-uttered audio and converts it into digital data.
[1694] A "means for capturing viewpoint images" is a device or method for acquiring images seen from the user's field of view in real time and recording them as digital data.
[1695] The "means for transmitting audio data and video data to a server" refers to a device or method for transferring the captured audio data and video data to a server via a communication network such as the Internet.
[1696] "Means for analyzing voice data and recognizing user intent" refers to algorithms or software that analyze received voice data and understand what the user is requesting.
[1697] "Means for analyzing video data and extracting necessary information" refers to algorithms or software that analyze received video data and identify and extract specific information (for example, text information or object information).
[1698] The "means for generating a response and transmitting result data to the terminal" refers to a device or method that generates an appropriate response based on the analyzed data and forwards the result to the user terminal.
[1699] "Means for notifying the user of the transmitted results" refers to a device or method for notifying the user of the transmitted results data at the terminal, including, for example, audio notification or visual display.
[1700] "Means for providing information corresponding to specific work instructions" refers to a method or device that combines audio and video capture to provide a user with specific instructions or information for a specific task or work.
[1701] A "pick list" is a list of instructions provided to workers for picking and sorting specific items at a distribution center.
[1702] "Inventory information" refers to data regarding the type, quantity, location, etc. of products stored in warehouses and logistics centers.
[1703] MODE FOR CARRYING OUT THE INVENTION
[1704] The present invention is a smart earphone system that provides real-time information to users in their daily lives. This system is particularly suitable for improving work efficiency in logistics centers. Specific embodiments for implementing this system are described below.
[1705] System configuration
[1706] This system consists of the following main components:
[1707] 1. User device (smart earphones and first-person camera)
[1708] 2. Server (AI processing unit)
[1709] 3. Internet connection
[1710] User terminal
[1711] The user terminal includes a microphone for capturing the user's voice and a first-person camera for capturing the user's point of view. The terminal also includes a communication module for transmitting data to a server.
[1712] server
[1713] The server includes an AI processing unit for analyzing the received audio and video data. This allows the server to recognize the user's intent and execute processing to provide the necessary information. The server also returns the resulting data to the device, which notifies the user of the results.
[1714] System Operation
[1715] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[1716] 1. Audio capture and transmission
[1717] The user speaks to the smart earphones, saying, "Tell me the next picklist." The device's microphone captures this voice and generates audio data. At the same time, a camera capturing the user's point of view captures the surroundings and generates video data. This data is temporarily stored in the device's local storage, compressed, and then sent to a server via the Internet.
[1718] 2. Analysis and processing on the server
[1719] The server analyzes the received voice data and recognizes the user's intent. For example, it extracts the intent "next pick list" from the voice. At the same time, it analyzes the video data and identifies shelf information, for example.
[1720] 3. Generating and returning result data
[1721] The server generates a picklist based on the analysis results, compresses this data again, and sends it to the terminal.
[1722] 4. Notice to Users
[1723] The terminal decompresses the received result data and notifies the user by voice, for example, by giving instructions such as "Please take product B from shelf A3."
[1724] Specific examples
[1725] Picklist Request Example
[1726] 1. User: "What's the next picklist?"
[1727] 2. Device: The microphone captures audio and the camera captures video of the surroundings. This data is sent to the server.
[1728] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the shelf information, then generates a pick list.
[1729] 4. Server: Sends the picklist to the terminal.
[1730] 5. Terminal: The picklist is announced to the user via voice: "Please pick up product B from shelf A3."
[1731] In this way, the system of the present invention can provide pick lists by voice instructions from logistics center workers and can analyze camera images to notify inventory information in real time. This system is expected to significantly improve work efficiency at logistics centers.
[1732] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1733] System program processing flow
[1734] Step 1:
[1735] The user speaks to the smart earphones and gives a voice command such as "Tell me the next picklist." This voice is captured by the microphone and generated as voice data.
[1736] Input: User's voice
[1737] Output: Audio data
[1738] How it works: The microphone in the smart earphones captures the user's voice and converts it into digital data.
[1739] Step 2:
[1740] The device's first-person camera captures images from the user's perspective and generates video data.
[1741] Input: Video from the user's point of view
[1742] Output: Video data
[1743] Specific operation: The device's camera captures video from the user's perspective in real time and stores it as digital data.
[1744] Step 3:
[1745] The captured audio and video data is temporarily stored in the terminal, compressed, and then transmitted to a server via the Internet.
[1746] Input: Audio data, video data
[1747] Output: Data sent to the server
[1748] Specific operation: Audio and video data is stored in the device's local storage, the data is compressed using a compression algorithm, and then sent to a server via an Internet connection.
[1749] Step 4:
[1750] The server analyzes the received voice data and recognizes the user's intent, for example, extracting the intent "next picklist."
[1751] Input: Audio data
[1752] Output: User intent (text format)
[1753] Specific operation: The server's AI processing unit uses a speech recognition engine (e.g., Google's speech recognition API) to convert the voice data into text and analyze the user's intent.
[1754] Step 5:
[1755] At the same time, the server analyzes the video data to identify, for example, shelf information.
[1756] Input: Video data
[1757] Output: Analysis results (shelf information)
[1758] Specific operation: Use the server's video analysis engine (such as OpenCV) to extract and identify specific information (e.g., shelf labels) from the video data.
[1759] Step 6:
[1760] The server generates a picklist based on the analysis results, compresses this data again, and sends it to the terminal.
[1761] Input: User intent, shelf information
[1762] Output: Picklist data
[1763] Specific operation: The generated picklist is compressed for transmission to the terminal, and then transmitted to the terminal via an Internet connection.
[1764] Step 7:
[1765] The terminal decompresses the received result data and notifies the user by voice, for example, by giving instructions such as "Please take product B from shelf A3."
[1766] Input: Compressed result data
[1767] Output: Audio notification
[1768] Specific operation: The device decompresses the compressed data and notifies the user by voice using a speech synthesis engine (e.g., Google TTS).
[1769] This series of processing steps allows workers at the logistics center to efficiently receive work instructions and respond immediately.
[1770] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1771] This invention is a smart earphone system that provides users with necessary information in real time in their daily lives and recognizes their emotions to provide appropriate feedback. This system captures audio and video from various viewpoints and analyzes the data on a server to generate a variety of responses. Furthermore, by combining it with an emotion engine, it is possible to provide adaptive services according to the user's emotions.
[1772] System configuration
[1773] This system consists of the following main components:
[1774] 1. User device (smart earphones and first-person camera)
[1775] 2. Server (AI processing unit and emotion engine)
[1776] 3. Internet connection
[1777] User terminal
[1778] The user terminal includes a microphone to capture the user's voice, a first-person camera to capture the user's point of view, and a communication module to transmit this data to a server via the Internet.
[1779] server
[1780] The server includes an AI processing unit for analyzing the received audio and video data. It is connected to an emotion engine, which is capable of recognizing the user's emotions from the received data. The server then generates appropriate information and responses based on the user's intentions and emotions, and notifies the user via the device.
[1781] System Operation
[1782] In an embodiment of the system, when a user speaks to the smart earbuds, the following occurs:
[1783] Audio capture and transmission
[1784] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and converts it into digital audio data. At the same time, a first-person camera captures video from the user's point of view and converts it into digital video data. This data is temporarily stored in local storage, then compressed and sent to a server using a secure protocol.
[1785] Server parsing and processing
[1786] The server analyzes the received voice data and converts it into text data using a speech recognition engine. For example, it recognizes the intent of "I want it translated." At the same time, it analyzes the video data, identifies the text area on the sign, and extracts the text information. It then uses an emotion engine to recognize the user's emotions from the voice and video data, generating information to indicate the urgency of the request and provide adaptive feedback. The server then translates the text information into the specified language using a translation engine.
[1787] Generate and return result data
[1788] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[1789] User Notification
[1790] The device receives and decompresses the resulting data. It then notifies the user of the translation results and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[1791] Specific examples
[1792] Translation request example
[1793] 1. User: "I want the text on the sign translated."
[1794] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[1795] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the characters on the sign. The emotion engine analyzes the user's urgency and emotion from the voice and video. The translation engine then translates the characters into the specified language.
[1796] 4. Server: Sends translation results and emotion recognition results to the device.
[1797] 5. Terminal: The user is notified of the translation result and emotional feedback via voice, such as "This sign says 'Business hours: 10:00-22:00'. Has your problem been resolved?"
[1798] Anomaly detection example
[1799] 1. User: (Unconscious, collapsed)
[1800] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[1801] 3. Server: Analyzes the video data and anomaly detection data and determines whether the user is in danger.
[1802] 4. Server: Generates and sends data to contact emergency services and provide them with the necessary information, as well as to notify registered emergency contacts.
[1803] 5. Terminal: The user is notified by voice that the emergency call has been completed. The user is told, "An ambulance has been dispatched. Please wait a moment."
[1804] In this way, the system of the present invention can provide information quickly and accurately in the user's daily life and provide adaptive feedback according to the user's emotions.
[1805] The processing flow will be explained below.
[1806] Specific processing steps
[1807] Example 1: Processing a translation request
[1808] Step 1:
[1809] The user speaks to the smart earphones and says, "I want the text on the sign translated."
[1810] Step 2:
[1811] The device's microphone captures the user's voice, which is then converted into a digital format.
[1812] Step 3:
[1813] The device's first-person camera captures video from the user's point of view, and the video data is converted into a digital format.
[1814] Step 4:
[1815] The audio and video data captured by the device is temporarily stored in local storage.
[1816] Step 5:
[1817] The terminal transmits the audio and video data to the server using a secure protocol.
[1818] Step 6:
[1819] The server receives the voice data and converts it into text data using a voice recognition engine. For example, a command such as "Please translate the text on the sign" is converted into text information.
[1820] Step 7:
[1821] The server receives the video data and uses computer vision technology to identify text regions in the video, then extracts the identified text information.
[1822] Step 8:
[1823] The server combines the results of speech recognition and video analysis to understand the user's intent. In this case, it recognizes the intent to "translate the text on the sign."
[1824] Step 9:
[1825] The emotion engine analyzes audio and video data to recognize the user's emotions. For example, it determines that the user is feeling anxious.
[1826] Step 10:
[1827] The server translates the text information into the specified language using a translation engine, for example, translating a Japanese sign into English.
[1828] Step 11:
[1829] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[1830] Step 12:
[1831] The terminal receives and decompresses the resulting data.
[1832] Step 13:
[1833] The device will then provide the user with a voice message with the result data, for example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[1834] Example 2: Emergency call due to abnormality detection
[1835] Step 1:
[1836] The user falls into an abnormal state such as losing consciousness.
[1837] Step 2:
[1838] The device's first-person camera captures the user's unusual situation.
[1839] Step 3:
[1840] The device acquires anomaly detection data from the biometric sensor and temporarily stores it in local storage.
[1841] Step 4:
[1842] The video data and anomaly detection data captured by the device are sent to the server using a secure protocol.
[1843] Step 5:
[1844] The server analyzes the received video data and anomaly detection data.
[1845] Step 6:
[1846] The server determines from the analysis results that the user is in danger.
[1847] Step 7:
[1848] The emotion engine analyzes the data and determines that the user is in an emergency.
[1849] Step 8:
[1850] The server automatically contacts the nearest emergency services and provides them with the necessary information, such as the user's current location and condition.
[1851] Step 9:
[1852] The server generates data to notify the registered emergency contacts and transmits this to the terminal.
[1853] Step 10:
[1854] The device receives the emergency call results and provides the information to relatives and friends via voice or other notification means.
[1855] Step 11:
[1856] The terminal notifies the user by voice that the ambulance has been dispatched, for example, by saying, "The ambulance has been dispatched. Please wait a moment."
[1857] The above are the specific processing steps in the system of the present invention.
[1858] Example 2
[1859] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1860] Conventional information provision systems have had difficulty efficiently analyzing the user's voice and visual perspective and providing appropriate information and feedback in real time. They also lacked the technology to recognize the user's emotions and generate appropriate responses accordingly. Furthermore, there was no mechanism for detecting abnormal situations and responding quickly, which created problems in ensuring the user's safety.
[1861] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing voice data and converting it into text data using a voice recognition engine, means for analyzing video data and extracting information using a video analysis engine, and means for recognizing the user's emotions from the voice data and video data using an emotion engine. This makes it possible to analyze the user's voice and video from their viewpoint in real time and provide adaptive information and feedback according to the user's intentions and emotions. Furthermore, in an emergency, abnormalities can be detected and emergency services can be quickly contacted, thereby ensuring the user's safety.
[1862] "User" refers to an individual who uses the system and receives information.
[1863] "Means for capturing audio" refers to a microphone and its peripheral devices for converting a user's speech into digital data.
[1864] The term "means for capturing an image of a viewpoint" refers to a camera and its peripheral devices for converting an image seen from the user's viewpoint into digital data.
[1865] "Means for transmitting captured audio and video data to a server" refers to a communications module for transmitting digital data to a server via the Internet or other network.
[1866] "Means for analyzing voice data and converting it into text data using a voice recognition engine" refers to voice recognition technology for processing voice data and converting it into text information.
[1867] "Means for analyzing video data and extracting information using a video analytics engine" refers to video analytics technology for processing video data to extract specific information.
[1868] "Means for recognizing a user's emotions from audio data and video data using an emotion engine" refers to a technology for analyzing a user's emotions by understanding the characteristics of audio and video.
[1869] "Means for generating a response based on the analyzed data and transmitting the resulting data to the terminal" refers to algorithms and communication modules for generating an appropriate response based on the analyzed information and transmitting the response to the terminal.
[1870] "Means for notifying the user of the transmitted results" refers to a speaker or display on the terminal for displaying or notifying the user of the results by voice.
[1871] "Text data" refers to the text information of a user's speech converted by a voice recognition engine.
[1872] "Means for extracting information" refers to the processes and algorithms used by the video analytics engine to extract specific information.
[1873] "Means for recognizing emotions" refers to a process for analyzing and identifying a user's emotions using an emotion engine.
[1874] "Results data" refers to data that includes response information generated by a server.
[1875] This invention relates to a system that provides users with necessary information in real time in their daily lives and recognizes their emotions to provide appropriate feedback. This system captures audio and video from various viewpoints and analyzes the data on a computer server, enabling it to generate a variety of responses. Furthermore, by combining it with an emotion engine, it becomes possible to provide adaptive services according to the user's emotions.
[1876] System configuration
[1877] This system consists of the following main components:
[1878] 1. User device (smart earphones and first-person camera)
[1879] 2. Server (AI processing unit and emotion engine)
[1880] 3. Internet connection
[1881] User terminal
[1882] The user terminal includes a microphone to capture the user's voice, a first-person camera to capture the user's point of view, and a communication module to transmit this data to a server via the Internet.
[1883] server
[1884] The server includes an AI processing unit for analyzing the received audio and video data. It is connected to an emotion engine, which is capable of recognizing the user's emotions from the received data. The server then generates appropriate information and responses based on the user's intentions and emotions, and notifies the user via the device.
[1885] Specific operation example
[1886] Audio capture and transmission
[1887] A user speaks to a smart earphone, saying, "I want the text on the sign translated." The device's microphone captures this voice and converts it into digital audio data. At the same time, a first-person camera captures video from the user's point of view and converts it into digital video data. This data is temporarily stored in local storage, then compressed and sent to a server using a secure protocol.
[1888] Server parsing and processing
[1889] The server analyzes the received voice data and converts it into text data using a speech recognition engine. For example, it recognizes the intent of "I want it translated." At the same time, it analyzes the video data, identifies the text area on the sign, and extracts the text information. It then uses an emotion engine to recognize the user's emotions from the voice and video data, generating information to indicate the urgency of the request and provide adaptive feedback. The server then translates the text information into the specified language using a translation engine.
[1890] Generate and return result data
[1891] The server generates result data including the translation result and the emotion recognition result, compresses it, and transmits it to the terminal.
[1892] User Notification
[1893] The device receives and decompresses the resulting data. It then notifies the user of the translation results and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "This sign says 'Business Hours: 10:00 - 22:00'. Is there anything I can help you with?"
[1894] Specific examples
[1895] Translation request example
[1896] 1. User: "I want the text on the sign translated."
[1897] 2. Terminal: The microphone captures the audio and the camera captures the image of the sign. These data are sent to the server.
[1898] 3. Server: The speech recognition engine converts the voice data into text, and the video analysis engine recognizes the characters on the sign. The emotion engine analyzes the user's urgency and emotion from the voice and video. The translation engine then translates the characters into the specified language.
[1899] 4. Server: Sends translation results and emotion recognition results to the device.
[1900] 5. Terminal: The user is notified of the translation result and emotional feedback via voice, such as "This sign says 'Business hours: 10:00-22:00'. Has your problem been resolved?"
[1901] Anomaly detection example
[1902] 1. User: (Unconscious, collapsed)
[1903] 2. Terminal: The camera captures the user's abnormal state and sends the abnormality detection data to the server.
[1904] 3. Server: Analyzes the video data and anomaly detection data and determines whether the user is in danger.
[1905] 4. Server: Generates and sends data to contact emergency services and provide them with the necessary information, as well as to notify registered emergency contacts.
[1906] 5. Terminal: The user is notified by voice that the emergency call has been completed. The user is told, "An ambulance has been dispatched. Please wait a moment."
[1907] In this way, the system of the present invention can provide information quickly and accurately in the user's daily life and provide adaptive feedback according to the user's emotions.
[1908] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1909] Step 1:
[1910] The user speaks to the smart earphones, saying, "I want the text on the sign to be translated." The user's voice becomes the input.
[1911] Step 2:
[1912] The device's microphone captures the user's voice and converts it from an analog audio signal into digital audio data, which is then output.
[1913] Step 3:
[1914] The device's first-person camera captures video from the user's point of view and converts the analog video signal into digital video data, which is then output.
[1915] Step 4:
[1916] The device temporarily stores digital audio and video data in local storage, which acts as a temporary buffer.
[1917] Step 5:
[1918] The device compresses the stored data (e.g., in gzip format) to make it easier to analyze, and the compressed data is output.
[1919] Step 6:
[1920] The device sends the compressed data to the server using a secure protocol (e.g., HTTPS), which then becomes the server's input.
[1921] Step 7:
[1922] The server analyzes the received voice data using a voice recognition engine (e.g., general voice recognition technology) and converts the digital voice data into text data, which is then output.
[1923] Step 8:
[1924] The server analyzes the text data and recognizes the user's intent (e.g., a request to "translate"). The recognition result becomes the input for the next step.
[1925] Step 9:
[1926] The server analyzes the received video data using a video analysis engine (e.g., general video recognition technology) to identify the text area on the sign. The identified text area becomes the output.
[1927] Step 10:
[1928] The server extracts text information from the specified text area, and this text information becomes the output.
[1929] Step 11:
[1930] The server uses an emotion engine (e.g., general emotion analysis technology) to recognize the user's emotion from the audio and video data. The emotion recognition result is the output.
[1931] Step 12:
[1932] The server generates a response based on the analysis results (user intent, text information, and emotion recognition results). The generated response becomes the output.
[1933] Step 13:
[1934] The server compresses the response data (e.g., in gzip format) and sends it to the terminal. The sent data becomes the terminal's input.
[1935] Step 14:
[1936] The device decompresses the received data and provides the translation results and feedback to the user in a voice message with a tone and content that adapts to the user's emotions, such as, "This sign says 'Business hours: 10:00-22:00'. Is there anything I can help you with?"
[1937] (Application example 2)
[1938] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1939] Conventional work support systems in factories often provide limited information and basic feedback, lacking adaptive support based on the worker's emotions and situation. They also lack early detection and rapid response to abnormal situations during work. This makes it difficult to improve work efficiency and ensure safety.
[1940] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1941] In this invention, the server includes a means for recognizing the user's emotions and generating an appropriate response based on the emotions, a means for transmitting work instructions and safety guidelines in real time within the factory, and a means for detecting abnormalities from received video data and automatically contacting emergency services. This enables adaptive support based on the emotions of workers, as well as early detection and rapid response to abnormal situations during work.
[1942] "User" refers to an individual who uses the system, particularly a worker who performs work within a factory.
[1943] "Audio data" refers to digital data that captures audio information uttered by a user.
[1944] "Visual data" refers to digital data that contains video information captured from a user's point of view.
[1945] "Server" refers to a computing device that analyzes the transmitted audio and visual data, extracts the necessary information, and generates a response for the user.
[1946] "Emotion recognition" refers to the process of analyzing a user's audio and visual data to identify the user's emotional state.
[1947] "Work instructions" refers to instruction information that conveys specific work content and procedures to users within a factory.
[1948] "Safety guidelines" refers to instruction documents that clearly state the safety measures and precautions that should be observed during work.
[1949] "Anomaly detection" refers to the process of analyzing received video data to identify situations that deviate from normal working conditions or emergencies.
[1950] "Emergency services" refers to external support services such as ambulances, fire departments, and police departments that respond to emergency situations during work.
[1951] This invention is a smart earphone system that provides real-time support to workers (users) as a work support system in factories. The system captures and analyzes the user's voice and visual data to recognize emotions and provide adaptive work instructions and safety guidelines. It also has the ability to detect abnormalities and automatically contact emergency services.
[1952] System configuration
[1953] This system consists of the following main components:
[1954] 1. User Device
[1955] The user terminal includes a smart earphone and a first-person camera worn by the worker. The terminal is equipped with a microphone for capturing the user's voice, a camera for capturing first-person perspective video, and a communication module for transmitting the data to a server.
[1956] 2. Server
[1957] The server is equipped with an AI processing unit and emotion engine for receiving and analyzing voice and visual data. It has the ability to analyze voice data to recognize the user's intentions and visual data to extract necessary information. The emotion engine also recognizes the user's emotions and generates adaptive responses according to those emotions.
[1958] System Operation
[1959] When a user speaks to the smart earbuds, the system works as follows:
[1960] 1. Audio and visual data capture
[1961] When a user requests a work instruction, the smart earphones' microphone captures audio and the camera captures video from the user's point of view, and these digital data are temporarily stored in local storage before being transmitted to a server using a secure protocol.
[1962] 2. Analysis and processing on the server
[1963] The server analyzes the received voice data and converts it into text data using a speech recognition engine. At the same time, it analyzes the visual data to extract the necessary information. For example, if a user requests, "Please tell me the next step in the process," the server recognizes the user's intention and generates information about the next step.
[1964] 3. Emotion Recognition and Response Generation
[1965] The emotion engine recognizes the user's emotions from audio and visual data, generating information to determine the urgency of the request and provide adaptive feedback. For example, if the user is feeling stressed, it will provide a safety warning.
[1966] 4. Generating and returning result data
[1967] The server generates result data including analysis results, emotion recognition results, and work instructions, compresses it, and sends it to the terminal.
[1968] 5. Notice to Users
[1969] The device receives and decompresses the resulting data. It then provides the user with work instructions and adaptive feedback in a voice message, with tone and content that reflects the user's emotions. For example, "Next, install part A. Safety first."
[1970] Specific examples
[1971] Work instruction example
[1972] The user speaks to the system, saying, "Please tell me the next steps," and the first-person camera captures the work environment. The system analyzes the audio and video, generates appropriate work instructions, and notifies the user by voice.
[1973] Example prompts to input to a generative AI model:
[1974] The user speaks to the smart earphones, "What are the next steps?", and the first-person camera captures the work environment. What steps can be taken to analyze the audio and video, generate appropriate work instructions, and provide voice prompts?
[1975] In this way, the system of the present invention can provide information quickly and accurately to workers in their daily work, and can provide adaptive feedback according to the user's emotions.
[1976] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1977] Step 1:
[1978] This is the step where the user speaks instructions to the smart earphone. For example, the user might say, "Please tell me the next step." This voice input triggers the start of the next processing step.
[1979] Step 2:
[1980] This is the step where the device captures the audio and converts it into digital audio data.
[1981] Input: User's voice
[1982] Output: Digital audio data
[1983] The device's microphone captures the user's voice, and an internal processor converts this voice into digital audio data.
[1984] Step 3:
[1985] The terminal simultaneously captures the visual data and converts it into digital video data.
[1986] Input: Video from the user's point of view
[1987] Output: Digital video data
[1988] A first-person camera mounted on the device captures video from the user's point of view and converts this video into digital data.
[1989] Step 4:
[1990] This is a step in which the terminal transmits the captured audio and video data to the server.
[1991] Input: Digital audio data, digital video data
[1992] Output: Send data to the server
[1993] A communication module in the terminal transmits the captured audio and video data to a server over a secure protocol.
[1994] Step 5:
[1995] This is the step where the server analyzes the received voice data and recognizes the user's intention.
[1996] Input: Digital audio data
[1997] Output: Text data and intent recognition results
[1998] The server's speech recognition engine converts the voice data into text data, analyzes it, and identifies the user's intent.
[1999] Step 6:
[2000] This is the step where the server analyzes the received visual data and extracts the necessary information.
[2001] Input: Digital video data
[2002] Output: Extracted information
[2003] The visual analytics engine analyzes the video data and identifies necessary information (such as the state of the work environment or specific objects).
[2004] Step 7:
[2005] This is the step where the server recognizes the user's emotions from the audio data and visual data.
[2006] Input: Digital audio data, digital video data
[2007] Output: User's emotional state
[2008] The emotion engine analyzes audio and video to identify the user's emotional state.
[2009] Step 8:
[2010] This is the step in which the server generates a response based on the analyzed data, creates result data, and sends it to the terminal.
[2011] Input: Intention recognition results, extracted information, emotional state
[2012] Output: Result data
[2013] The server generates an appropriate response based on the user's intention, information about the working environment, and emotional state, and creates result data, which is then compressed and sent to the terminal.
[2014] Step 9:
[2015] This is the step in which the terminal decompresses the received result data and notifies the user.
[2016] Input: Compressed result data
[2017] Output: Audio notification
[2018] The terminal decompresses the received result data and notifies the user by voice of appropriate work instructions and precautions, such as "Next, please install part A. Safety first."
[2019] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2020] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2021] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2022] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2023] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2024] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2025] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2026] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2027] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2028] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2029] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2030] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2031] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2032] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2033] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2034] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2035] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2036] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2037] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2038] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2039] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2040] The following is further disclosed regarding the above embodiment.
[2041] (Claim 1)
[2042] means for capturing the user's voice;
[2043] a means for capturing video of a user's point of view;
[2044] means for transmitting the captured audio and video data to a server;
[2045] A means for analyzing the voice data by a server and recognizing the user's intention;
[2046] A means for analyzing the video data by a server and extracting necessary information;
[2047] means for generating a response based on the analyzed data and transmitting the result data to the terminal;
[2048] means for notifying the user of the transmitted result;
[2049] A system including:
[2050] (Claim 2)
[2051] 2. The system according to claim 1, which recognizes the user's intention by analyzing voice data, and extracts and translates text information by analyzing video data.
[2052] (Claim 3)
[2053] 10. The system of claim 1, wherein the system detects an abnormality from the received video data and automatically contacts emergency services.
[2054] "Example 1"
[2055] (Claim 1)
[2056] means for capturing the user's voice;
[2057] a means for capturing video of a user's point of view;
[2058] means for compressing the captured audio and video data;
[2059] means for transmitting the compressed audio data and video data to a server;
[2060] A means for analyzing the voice data by using a voice recognition engine in a server and recognizing the user's intention;
[2061] A server analyzes the video data using an image analysis means and extracts necessary information;
[2062] means for translating the extracted information into a specified language using a translation engine;
[2063] means for generating a response based on the analyzed and translated data and transmitting the resulting data to the terminal;
[2064] means for notifying the user of the transmitted result using a speech synthesis engine;
[2065] A system including:
[2066] (Claim 2)
[2067] 2. The system according to claim 1, wherein the system recognizes a user's intention by analyzing voice data, extracts text information by analyzing video data, and translates the extracted text information into a specified language.
[2068] (Claim 3)
[2069] 10. The system of claim 1, wherein the system detects anomalies in the received video data and automatically contacts emergency services.
[2070] "Application Example 1"
[2071] Rewritten claims
[2072] (Claim 1)
[2073] means for capturing the user's voice;
[2074] a means for capturing video of a user's point of view;
[2075] means for transmitting the captured audio and video data to a server;
[2076] A means for analyzing the voice data by a server and recognizing the user's intention;
[2077] A means for analyzing the video data by a server and extracting necessary information;
[2078] means for generating a response based on the analyzed data and transmitting the result data to the terminal;
[2079] means for notifying the user of the transmitted result;
[2080] a means for combining audio and video capture to provide information corresponding to a particular work instruction;
[2081] A system including:
[2082] (Claim 2)
[2083] 2. The system according to claim 1, which recognizes the user's intention by analyzing voice data, and extracts and translates text information by analyzing video data.
[2084] (Claim 3)
[2085] 10. The system of claim 1, wherein the system detects an abnormality from the received video data and automatically contacts emergency services.
[2086] (Claim 4)
[2087] The system according to claim 1, wherein the system provides a pick list in response to user voice instructions at a logistics center, analyzes camera footage, and notifies inventory information in real time.
[2088] "Example 2: Combining Emotion Engines"
[2089] (Claim 1)
[2090] means for capturing the user's voice;
[2091] a means for capturing video of a user's point of view;
[2092] means for transmitting the captured audio and video data to a computer server;
[2093] A means for analyzing the voice data and converting it into text data using a voice recognition engine;
[2094] means for analyzing the video data and extracting information using a video analytics engine;
[2095] means for recognizing a user's emotion from audio data and video data using an emotion engine;
[2096] means for generating a response based on the analyzed data and transmitting the resulting data to the terminal;
[2097] means for notifying the user of the transmitted result;
[2098] A system including:
[2099] (Claim 2)
[2100] 2. The system according to claim 1, which recognizes a user's intention by analyzing voice data, extracts text information by analyzing video data, and translates it into a specified language.
[2101] (Claim 3)
[2102] 10. The system of claim 1, wherein the system detects anomalies in the received video data and automatically contacts emergency services.
[2103] "Application example 2 when combining emotion engines"
[2104] (Claim 1)
[2105] means for capturing the user's voice;
[2106] means for capturing visual data of a user's viewpoint;
[2107] means for transmitting the captured audio and visual data to a server;
[2108] A means for analyzing the voice data by a server and recognizing the user's intention;
[2109] A server analyzes the visual data and extracts necessary information;
[2110] means for generating a response based on the analyzed data and transmitting the result data to the terminal;
[2111] means for notifying the user of the transmitted result;
[2112] means for recognizing a user's emotion and generating an appropriate response in response to the emotion;
[2113] A means of communicating work instructions and safety guidelines in real time within the factory,
[2114] A system including:
[2115] (Claim 2)
[2116] 2. The system according to claim 1, which recognizes the user's intention by analyzing voice data, and extracts and translates text information by analyzing video data.
[2117] (Claim 3)
[2118] 10. The system of claim 1, wherein the system detects anomalies in the received video data and automatically contacts emergency services. [Explanation of symbols]
[2119] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for capturing the user's voice; a means for capturing video of a user's point of view; means for transmitting the captured audio and video data to a server; A means for analyzing the voice data by a server and recognizing the user's intention; A means for analyzing the video data by a server and extracting necessary information; means for generating a response based on the analyzed data and transmitting the result data to the terminal; means for notifying the user of the transmitted result; A system including:
2. 2. The system according to claim 1, wherein the system recognizes the user's intention by analyzing the voice data, and extracts and translates text information by analyzing the video data.
3. 2. The system according to claim 1, wherein an abnormality is detected from the received video data and an emergency service is automatically contacted.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A