System
The smart glasses-based system addresses the inefficiency of manual task ticketing by automatically capturing and analyzing video and audio data to generate and manage task tickets, enhancing productivity and reducing errors.
Patent Information
- Application Number
- JP2024119077
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Modern office and factory workers face the challenge of manually creating task tickets in task management systems, which is time-consuming, prone to errors, and inefficient in capturing tasks arising during meetings and conversations.
A system using smart glasses with integrated cameras and microphones collects real-time video and audio data, transmitting it to a server for conversion to text and extraction of task information using image and speech recognition, followed by automatic generation and registration of task tickets in a management system.
This system allows for efficient, error-free management of tasks without manual intervention, reducing the burden and improving productivity by automatically generating and notifying users of task tickets.
Smart Images

Figure 2026018016000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Modern office workers are often asked to complete tasks via meetings, chat, email, and various other means. In order to efficiently manage these tasks without missing any, it is necessary to manually create tickets in a task management system, but this process is time-consuming and labor-intensive. Manual entry also makes it prone to errors and omissions. Given this situation, there is a demand for a system that can automatically understand and manage tasks without bothering the user. [Means for solving the problem]
[0005] This invention proposes a system that uses a camera and microphone built into smart glasses to collect video and audio data from the user's perspective in real time and transmits that data to a server via the Internet. The server converts the audio data into text using voice recognition technology and then extracts task information from the video data using image recognition technology. Natural language processing technology analyzes the task information from this data and automatically generates task tickets. The generated task tickets are registered in a task management system and notified to the user. This allows users to manage tasks hassle-free, saving time and effort.
[0006] An "integrated camera" is a camera device built into smart glasses to capture images from the user's point of view in real time.
[0007] "Built-in microphone" refers to a microphone device built into smart glasses to collect voice data in real time.
[0008] "Video data" refers to visual information data acquired using a built-in camera.
[0009] "Audio data" refers to sound information collected using the built-in microphone.
[0010] "Server" is a central computer system that receives and analyzes collected video and audio data.
[0011] The "transmission means" is a means for transmitting the collected video data and audio data to the server.
[0012] "Speech recognition technology" is a technology that converts voice data into text data.
[0013] "Character data" refers to data in text format that has been converted using voice recognition technology.
[0014] "Image recognition technology" is a technology that extracts task information from video data.
[0015] "Task information" is information such as specific tasks and instructions that a user should perform.
[0016] "Natural language processing technology" is a technology for analyzing task information from text data.
[0017] A "task ticket" is a work instruction sheet that is generated based on task information and registered in the task management system.
[0018] A "task management system" is a system that manages task tickets and allows users to understand the task status.
[0019] The "notification means" is a means for notifying the user of the generated task ticket. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be explained.
[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0028] [First embodiment]
[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0041] To implement this invention, we will build a system using smart glasses and a server connected to them. This system collects video and audio data from the user's point of view, analyzes it on the server, generates task tickets, and registers them in a task management system.
[0042] Program processing overview
[0043] Data collection
[0044] The device uses the smart glasses' built-in camera to capture real-time video from the user's perspective, and a built-in microphone to collect the user's conversation and surrounding sounds, allowing for accurate recording of task information during meetings and conversations.
[0045] Example: The device simultaneously records whiteboard content and conversations during meetings. When the user is looking at a PC screen, the device also captures on-screen task information.
[0046] Data transmission
[0047] The devices transmit the collected video and audio data in real time to a server over an internet connection, where the data is compressed for efficient transmission and encrypted to ensure the information is secure.
[0048] Example: A device uses a Wi-Fi connection to send video and audio data to a server. Data can be sent even when you are away from home.
[0049] Data analysis
[0050] The server inputs the received video data into an image recognition engine to extract important elements including task information. At the same time, it inputs the audio data into a voice recognition engine and converts it into text data. Then, using natural language processing technology, it analyzes the task information from this text data.
[0051] Example: The server converts audio data such as "Write a report and submit it by next Monday" into text and extracts the task elements "Write a report," "Submit," and "Next Monday." It also analyzes a whiteboard note from video data that reads "Deadline: May 20th" and recognizes the deadline.
[0052] Generate a Task Ticket
[0053] The server automatically generates a task ticket based on the parsed task information, which includes information such as the task summary, details, deadline, and assignee, allowing users to manage tasks accurately and quickly without any manual intervention.
[0054] Example: A task ticket with the following name will be automatically generated: "Create a report", "Submission deadline: next Monday", "Responsible person: User".
[0055] Ticket notifications and management
[0056] The server registers the generated task ticket in the task management system and has a notification means for notifying the user. The user receives the notification, accesses the task management system, checks the generated task ticket, and can correct or enter additional information as necessary.
[0057] Example: The server notifies the user that a "New Task Ticket: Create Report" has been created, and the user opens a task management app to check the contents of the ticket. If necessary, they can enter notes or supplementary information and manage the progress of the task.
[0058] This system allows users to efficiently manage various tasks that arise during meetings, conversations, and work, significantly reducing the burden and errors of manual ticket creation.
[0059] The processing flow will be explained below.
[0060] Step 1:
[0061] The device uses a built-in camera to capture video from the user's point of view in real time, including the contents of the whiteboard, presentation slides, and PC screen during the meeting.
[0062] Step 2:
[0063] The device uses a built-in microphone to collect audio data in real time, recording the user's conversations and the sounds of the surrounding area.
[0064] Step 3:
[0065] The devices transmit the collected video and audio data to a server over the Internet, where the data is pre-compressed and encrypted to ensure efficient and secure transmission.
[0066] Step 4:
[0067] The server inputs the received voice data into a speech recognition engine and converts the voice into text data, making the conversation available in text format.
[0068] Step 5:
[0069] The server inputs the received video data into an image recognition engine and extracts task-related information from the video, for example, identifying text written on a whiteboard or task information displayed on a PC screen.
[0070] Step 6:
[0071] The server uses natural language processing technology to analyze the converted text data and identify task information, such as the task content, deadline, and person in charge, from the conversation and video data.
[0072] Step 7:
[0073] The server automatically generates a task ticket based on the analyzed task information. The generated task ticket includes information such as the task summary, details, deadline, and person in charge.
[0074] Step 8:
[0075] The server registers the generated task ticket in the task management system, which allows tasks to be centrally managed and users to check them later.
[0076] Step 9:
[0077] The server will notify the user of the newly created task ticket, via email, push notification, or other methods.
[0078] Step 10:
[0079] The user receives a notification and accesses the task management system to view the generated task ticket, enter additional notes and details as needed, and manage the task's progress.
[0080] Example 1
[0081] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0082] Conventional task management systems require task information to be entered manually, a time-consuming process that poses a high risk of input errors. Furthermore, there was no way to efficiently manage tasks that arise during meetings or conversations, which meant there was a risk of important information being overlooked or tasks being left out.
[0083] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0084] In this invention, the server includes means for capturing video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for compressing and encrypting the transmitted data, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, and means for notifying the user of the generated task tickets. This makes it possible to reliably collect and analyze task information in real time even during meetings and conversations, and automatically generate and manage task tickets.
[0085] An "integrated camera" is a camera that is built into a device and has the ability to capture video from the user's perspective in real time.
[0086] An "internal microphone" is a microphone that is built into a device and has the ability to collect audio data in real time.
[0087] A "server" is a computer system on a network that receives collected data and performs analysis and task ticket generation.
[0088] "Real-time" is a time processing concept that means data is generated, transmitted, and analyzed almost immediately.
[0089] "Speech recognition technology" is a technology that analyzes voice data and converts the content into text data.
[0090] "Character data" is text information obtained from voice data using voice recognition technology.
[0091] "Image recognition technology" is a technology that extracts specific information and patterns from video data.
[0092] "Natural language processing technology" is a general term for algorithms and technologies that analyze text data and understand, classify, and extract its content.
[0093] A "task ticket" is a manageable unit of a task that includes information such as task details, deadlines, and people in charge.
[0094] "Task Management System" means software or platform that registers and tracks generated task tickets and assists users in managing their tasks.
[0095] "Compression" is the process of reducing the size of data, a technique used to improve transmission efficiency.
[0096] "Encryption" is a technology that converts data to protect it and prevent unauthorized access or tampering during transmission.
[0097] A "push notification" is a notification message sent in real time from a server to a user's device.
[0098] To implement this invention, it is necessary to build a system using smart glasses and a server connected to them. The present invention collects video and audio data from the user's perspective, analyzes it on the server, generates task tickets, and registers them in a task management system.
[0099] First, the user puts on the smart glasses. The smart glasses are equipped with a built-in camera and microphone, which are used to collect video and audio data in real time. Specifically, the camera captures the video the user is looking at, and the microphone collects the conversations of the user and those around them. Typical smart glasses products can be used as the specific hardware for this purpose.
[0100] The device then transmits the collected video and audio data in real time to a server. During this transmission process, the data is compressed and encrypted. An internet connection is used as the communication medium. At this stage, data security and efficient transfer are crucial.
[0101] The server inputs the received video data into an image recognition engine and extracts key elements, including task information. Software used can include Google Cloud Vision API for image recognition, Google Cloud Speech-to-Text for voice recognition, and NLP technologies such as spaCy and Transformers for natural language processing. For example, if voice data such as "Write a report and submit it by next Monday" is sent to the server, it is converted into text using voice recognition technology, and task elements such as "Write a report," "Submit," and "Next Monday" are extracted. Furthermore, a whiteboard note from the video data stating "Deadline: May 20th" is analyzed, and the deadline is recognized.
[0102] Next, the server automatically generates a task ticket based on the analyzed task information. This task ticket includes the task summary, details, deadline, person in charge, etc. For example, specific information such as "Create a report," "Submission deadline: next Monday," and "Person in charge: User" is written in the task ticket.
[0103] The generated task ticket is registered in the task management system by the server. Commonly used platforms such as Trello and Jira can be used as the task management system. The server notifies the user of the generated task ticket. Possible notification methods include push notification and email.
[0104] Finally, the user accesses the task management system based on the received notification and checks the generated task ticket. The user can enter corrections or additional information into the task ticket as necessary to manage the progress of the task.
[0105] Specific examples
[0106] For example, if a user is told during a meeting to "write a report and submit it by next Monday," the user's video and audio will be collected through smart glasses. The device will then send the data to a server, which will then use speech recognition and natural language processing to extract task information. This will automatically generate a task ticket in the task management system with the following information: "Create a report," "Submission deadline: next Monday," and "Responsible party: User."
[0107] Prompt Sentence Examples
[0108] "Interpret the notes on the whiteboard during the meeting and extract important task information. For example, generate a task ticket with the content 'Submit report due May 20th'."
[0109] This system makes it possible to efficiently manage the various tasks that arise during meetings and conversations, significantly reducing the effort required to manually enter tasks and the risk of entering them incorrectly.
[0110] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0111] Step 1:
[0112] The device uses the camera built into the smart glasses to capture video from the user's point of view in real time, while simultaneously using the built-in microphone to collect the user's conversation and surrounding sounds. The input data is video data and audio data, and the output data is the collected raw data.
[0113] Specific behavior:
[0114] The device activates the smart glasses camera.
[0115] The device begins collecting audio data using the built-in microphone.
[0116] When a user is looking at notes on a whiteboard during a meeting, the terminal records the contents of the whiteboard as video.
[0117] When a user speaks, the microphone in the smart glasses picks up the sound.
[0118] Step 2:
[0119] The terminal transmits collected video and audio data to the server in real time. The data is compressed for efficiency and encrypted for security. The input data is the collected raw data. The output data is the compressed and encrypted data.
[0120] Specific behavior:
[0121] The device compresses the video and audio data.
[0122] The device encrypts the data.
[0123] The device sends the data to the server through an internet connection.
[0124] The server decompresses and decrypts the received data.
[0125] Step 3:
[0126] The server inputs the received video data into an image recognition engine to extract important elements including task information. At the same time, it inputs the audio data into a speech recognition engine and converts it into text data. The input data is raw data that has been decompressed and decrypted. The output data is the analyzed task information.
[0127] Specific behavior:
[0128] The server sends the video data to an image recognition engine, for example, using the Google Cloud Vision API.
[0129] The server sends the voice data to a speech recognition engine and converts it into text data. For example, Google Cloud Speech-to-Text is used.
[0130] The server sends the converted text data to a natural language processing engine, which analyzes and extracts task information. For example, we use spaCy.
[0131] The server extracts task details such as "write a report," "submit," and "next Monday" from the voice data.
[0132] The server analyzes the whiteboard note "Deadline: May 20th" from the video data and recognizes the task deadline.
[0133] Step 4:
[0134] The server automatically generates a task ticket based on the parsed task information. The task ticket contains information such as the task summary, details, deadline, and person in charge. The input data is the parsed task information. The output data is the generated task ticket.
[0135] Specific behavior:
[0136] The server automatically generates a task ticket based on the analyzed task information.
[0137] The server writes information such as "Create report," "Submission deadline: next Monday," and "Responsible person: User" into the task ticket.
[0138] Step 5:
[0139] The server registers the generated task ticket in the task management system and notifies the user. The input data is the generated task ticket. The output data is the notified task ticket information.
[0140] Specific behavior:
[0141] The server registers the task ticket in a task management system, such as Trello or Jira.
[0142] The server notifies the user of the new task ticket creation via push notification or email.
[0143] A user accesses the task management system to view the new task ticket and enter notes or additional information as needed.
[0144] (Application example 1)
[0145] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0146] In modern factory environments, many tasks are still performed manually, and efficiency and accuracy are essential. However, workers spend a lot of time manually creating and managing task tickets, resulting in reduced productivity. Furthermore, there is no system in place to accurately record the various tasks that occur during work and instantly notify robots. This disrupts the smooth progress of work within the factory, making it difficult to address any errors or delays that arise.
[0147] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0148] In this invention, the server includes means for capturing video from a user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, means for notifying the generated task tickets to users, means for having workers wear smart glasses in a factory environment to collect task details, send the collected data to the server for analysis, and initiate task notification to factory robots, and means for supporting workers to receive work instructions in real time and perform tasks efficiently, thereby enabling automation and improved accuracy of task management within the factory.
[0149] "Built-in camera" refers to a camera module built into an apparatus or device that captures video from the user's perspective in real time.
[0150] "Built-in microphone" refers to a microphone built into an apparatus or device for collecting audio data in real time.
[0151] A "server" refers to a computer system that processes data over a network and has storage and analysis capabilities.
[0152] "Speech recognition technology" refers to technology that analyzes voice data and converts that voice into text data.
[0153] "Image recognition technology" refers to technology for analyzing video data and extracting specific information from it.
[0154] "Natural language processing technology" refers to technology that analyzes text data and extracts and understands information based on human language.
[0155] A "task ticket" is an electronic card or format that compiles information about a specific task, including the task name, details, deadline, and assignee.
[0156] A "task management system" refers to software for registering and managing generated task tickets.
[0157] "Smart glasses" refers to a wearable eyeglass-type device that has augmented reality and information display functions.
[0158] "Factory robot" refers to a robot used to automate production tasks in a factory.
[0159] "Worker" refers to a person who actually performs manual work or supervises in a factory or production site.
[0160] To implement this invention, a system using a server linked to smart glasses is constructed as follows.
[0161] Data collection
[0162] Device:
[0163] The user wears the smart glasses, and while working, the built-in camera captures video from the user's point of view in real time. The built-in microphone collects audio data. Examples of hardware used to collect this data include the built-in camera from Logitech and the microphone from Blue Yeti.
[0164] Specific examples
[0165] Workers on the factory assembly line wear smart glasses and perform their tasks while receiving verbal instructions on part installation procedures and important points to note, which then record the work procedures as visual and audio data.
[0166] Data transmission
[0167] Device:
[0168] The collected video and audio data is sent to a server via the Internet, where it is compressed and encrypted to ensure efficient and secure transmission.
[0169] Data analysis
[0170] server:
[0171] The server inputs the received video data into an image recognition engine (e.g., OpenCV, TensorFlow) to extract key elements including task information. It also inputs the audio data into a speech recognition engine (e.g., Google Speech-to-Text API) to convert it into text data. It then uses natural language processing technology (e.g., NLTK, spaCy) to analyze the text data for specific task information.
[0172] Specific examples
[0173] The server converts the voice data, "Tighten this bolt according to step 7," into text and extracts the task elements "bolt tightening task" and "step 7." It also analyzes and recognizes signs from the video data that include the instruction "inspect after task completion."
[0174] Task ticket generation and notification
[0175] server:
[0176] A task ticket is automatically generated based on the analyzed task information. The task ticket contains information such as the task summary, details, deadline, and person in charge. The generated task ticket is registered in the task management system, and its contents are notified to the user.
[0177] Specific examples
[0178] Task tickets for "Bolt tightening," "Step 7," and "Inspect after completion" are automatically generated and displayed on the worker's smart glasses. The same information is also sent to the factory robot, which automatically supports the necessary tasks.
[0179] Real-time support for work instructions
[0180] Terminals and robots:
[0181] The smart glasses help workers receive real-time work instructions to carry out their work efficiently. They also display the next steps and important points to note, and send instructions to factory robots.
[0182] Prompt Sentence Examples
[0183] If you're having a discussion in a conference room:
[0184] I need to write and submit the report by next Monday.
[0185] For factory assembly lines:
[0186] "When you assemble this part to the next block, you need to tighten the nuts."
[0187] These processes significantly improve task management and work efficiency within the factory, reducing the burden and errors of manual ticket creation.
[0188] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0189] Program processing flow
[0190] Step 1: Data collection
[0191] Device:
[0192] Using the built-in camera, video data is captured in real time from the user's perspective.
[0193] The built-in microphone is used to collect the user's voice data in real time.
[0194] Input: Video data, audio data
[0195] Output: Collected video and audio files
[0196] Step 2: Send data
[0197] Device:
[0198] The collected video and audio data is sent to a server via the Internet.
[0199] Data is compressed and encrypted for secure and efficient transfer.
[0200] Input: Collected video and audio files
[0201] Output: Data sent to the server
[0202] Step 3: Data analysis (audio)
[0203] server:
[0204] The voice data is input into a voice recognition engine, and the voice is converted into text data.
[0205] Use voice recognition technology (e.g., Google Speech-to-Text API).
[0206] Input: Audio data
[0207] Output: Text data
[0208] Step 4: Data analysis (video)
[0209] server:
[0210] The video data is input into an image recognition engine, and task information is extracted from the video.
[0211] Uses image recognition technology (e.g., OpenCV, TensorFlow).
[0212] Input: Video data
[0213] Output: Extracted task information
[0214] Step 5: Parse task information
[0215] server:
[0216] Text data is input into natural language processing technology to analyze task information.
[0217] Use natural language processing technology (e.g., NLTK, spaCy).
[0218] Input: Text data
[0219] Output: Parsed task information
[0220] Step 6: Generate a task ticket
[0221] server:
[0222] Based on the parsed task information, a task ticket is automatically generated.
[0223] The task information includes a summary of the task, details, deadlines, and the person in charge.
[0224] Input: Parsed task information
[0225] Output: Generated task ticket
[0226] Step 7: Task ticket registration and notification
[0227] server:
[0228] The generated task ticket is registered in the task management system and notified to the user.
[0229] Input: Generated task ticket
[0230] Output: Notify user and register with task management system
[0231] Step 8: Real-time support for work instructions
[0232] Terminals and robots:
[0233] Workers receive work instructions in real time and are supported to carry out their work efficiently.
[0234] The smart glasses display the next steps and precautions to take, and send instructions to factory robots.
[0235] Input: Task ticket information
[0236] Output: Work instructions for smart glasses and factory robots
[0237] Specific operation explanation
[0238] Step 1: The device uses the smart glasses' built-in camera and microphone to collect real-time video and audio from the worker's perspective. The worker explains how to install the parts as they go along, collecting the necessary data.
[0239] Step 2: The collected video and audio data is sent to a server via the internet. The data is compressed and encrypted for efficient and secure transmission.
[0240] Step 3: The server uses a speech recognition engine to convert the voice data into text data. For example, the voice saying "Tighten this bolt according to step 7" is converted into text data.
[0241] Step 4: The server uses a video recognition engine to extract task information from the video data. For example, it can read "Step 7" and "Check after inspection" from signs and labels in the video.
[0242] Step 5: The server uses natural language processing technology to analyze specific task information from the converted text data. For example, it extracts task elements such as "bolt tightening," "step 7," and "inspection after completion of work."
[0243] Step 6: The server automatically generates a task ticket based on the parsed task information. The task ticket includes the task name, details, deadline, and person in charge.
[0244] Step 7: The generated task ticket is registered in the task management system and simultaneously notified to the worker's smart glasses and factory robots. The worker can then view the new task through the smart glasses.
[0245] Step 8: Using smart glasses, workers receive real-time instructions on the next steps and important points to note, allowing them to work efficiently. Factory robots also receive instructions and automatically assist with the designated tasks.
[0246] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0247] To implement this invention, we will build a system using smart glasses and a server equipped with an emotion engine linked to the glasses. This system collects video and audio data from the user's perspective, analyzes it on the server, generates task tickets, performs emotion recognition, registers them in a task management system, and notifies the user.
[0248] Program processing overview
[0249] Data collection
[0250] The device uses the smart glasses' built-in camera to capture real-time video from the user's point of view, and a built-in microphone to collect the user's conversation and surrounding sounds. The emotion engine uses this data to analyze the user's emotions.
[0251] Example: The device simultaneously records the content of the whiteboard and conversations during a meeting. When the user is looking at the PC screen, the device also captures the task information on the screen. Furthermore, the device analyzes the user's emotions from facial expressions and tone of voice while they are speaking.
[0252] Data transmission
[0253] The device transmits the collected video, audio, and emotion analysis data to a server via the internet. This transmission uses an internet connection. The data is compressed to improve transmission efficiency and encrypted to ensure the security of the information.
[0254] Example: A device uses a Wi-Fi connection to send video and audio data, as well as emotion analysis data, to a server. Data can be sent even when the user is away from home.
[0255] Data analysis
[0256] The server inputs the received voice data into a voice recognition engine, which converts the voice into text data. It then inputs the video data into an image recognition engine, which extracts task-related information from the video. It then uses natural language processing technology to analyze the task information from the text data. The emotion engine also analyzes the user's emotions from the voice and video data.
[0257] Example: The server converts voice data such as "Write a report and submit it by next Monday" into text and extracts the task elements "Write a report," "Submit," and "Next Monday." It also analyzes a whiteboard note from video data that reads "Deadline: May 20th" to recognize the deadline. At the same time, the emotion engine analyzes whether the user's voice sounds tense or whether they are smiling.
[0258] Generate a Task Ticket
[0259] The server automatically generates a task ticket based on the parsed task information, which includes the task summary, details, deadline, assignee, and user sentiment information, allowing for more context-sensitive responses.
[0260] Example: A task ticket containing "Write a report", "Submission deadline: next Monday", "Responsible party: User", and "User emotion: Tension" is automatically generated.
[0261] Ticket notifications and management
[0262] The server registers the generated task ticket in the task management system and has a notification means to notify the user. The user receives the notification, accesses the task management system to check the generated task ticket, and can correct or add additional information as necessary. It is also possible to adjust the priority and response method of the task based on the emotional information.
[0263] Example: The server notifies the user that a "New Task Ticket: Create Report" has been created, and the user opens a task management app to check the contents of the ticket. If necessary, the user can enter notes or supplementary information and manage the progress of the task. Taking into account the user's emotional information, if the user is feeling nervous, the system will take appropriate measures such as lowering the priority of the task.
[0264] This system allows users to efficiently manage various tasks that arise during meetings, conversations, and work, and enables detailed task responses that take into account the user's emotional state, significantly reducing the burden and errors of manual task entry and improving users' overall work efficiency and satisfaction.
[0265] The processing flow will be explained below.
[0266] Step 1:
[0267] The device uses a built-in camera to capture real-time video from the user's point of view, including the contents of the whiteboard, presentation slides, and PC screen during the meeting.
[0268] Step 2:
[0269] The device uses a built-in microphone to collect voice data in real time, recording the user's conversation and surrounding sounds, and an emotion engine analyzes the tone and tempo of the voice to recognize the user's emotions.
[0270] Step 3:
[0271] The device then sends the collected video, audio, and emotion analysis data to a server via the internet. The data is compressed to improve transmission efficiency and encrypted to ensure the security of the information.
[0272] Step 4:
[0273] The server inputs the received voice data into a speech recognition engine and converts the voice into text data, making the conversation available in text format.
[0274] Step 5:
[0275] The server inputs the received video data into an image recognition engine and extracts task-related information from the video, for example, identifying text written on a whiteboard or task information displayed on a PC screen.
[0276] Step 6:
[0277] The server uses natural language processing technology to analyze task information from the converted text data. Specifically, it identifies the task content, deadline, and person in charge from the conversation and video data. In addition, an emotion engine analyzes the user's emotions from the audio and video data.
[0278] Step 7:
[0279] The server automatically generates a task ticket based on the analyzed task information and the user's emotion information. The generated task ticket includes the task summary, details, deadline, assignee, and user's emotion information.
[0280] Step 8:
[0281] The server registers the generated task ticket in the task management system. At that time, it is also possible to set the priority of the task based on the user's emotional information.
[0282] Step 9:
[0283] The server will notify the user of the newly created task ticket, via email, push notification, or other methods.
[0284] Step 10:
[0285] Users receive a notification and access their task management system to view the generated task ticket, enter additional notes and details as needed, manage the task's progress, and even adjust the task's priority based on sentiment information.
[0286] Example 2
[0287] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0288] In today's business environment, where a variety of tasks are being carried out simultaneously, manually inputting and managing tasks by users is laborious and prone to errors. Furthermore, task management that takes into account the user's emotional state is difficult, resulting in reduced work efficiency and satisfaction. Current task management systems are unable to automatically generate tasks in real time or manage tasks that take into account emotional information, so improvements are needed to address this issue.
[0289] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for compressing and encrypting the collected video data and audio data and transmitting them to the server, means for converting audio data into text data using voice recognition technology, means for extracting task information from video data using image recognition technology, means for analyzing task information from text data using natural language processing technology, means for analyzing user emotions from the collected video data and audio data using emotion analysis technology, means for automatically generating task tickets based on task information, means for adding emotion information to the generated task tickets, means for registering the generated task tickets in a task management system, and means for notifying the user of the generated task tickets. This enables efficient and accurate task management that also takes the user's emotional state into consideration.
[0290] A "built-in camera" is a photographing device built into a terminal to capture images from the user's point of view in real time.
[0291] A "built-in microphone" is a voice collection device built into a terminal for collecting voice data in real time.
[0292] "Collected Video Data" means visual information captured in real time by the built-in camera.
[0293] "Collected Audio Data" means audio information collected in real time by the built-in microphone.
[0294] "Means of compressing, encrypting and transmitting" refers to the technical means of compressing collected data for efficient transmission, encrypting it to ensure the security of communications and transmitting it to a server via the Internet.
[0295] "Speech recognition technology" is a technology that analyzes collected voice data and converts it into text data.
[0296] "Character data" is voice information analyzed by voice recognition technology expressed in text format.
[0297] "Image recognition technology" is a technology that analyzes collected video data and extracts task information from the video.
[0298] "Natural language processing technology" is a technology that analyzes task information from text data and understands meaning and intent.
[0299] "Emotion analysis technology" is a technology that analyzes collected video and audio data to recognize the user's emotional state.
[0300] "Task information" refers to information related to tasks and instructions extracted from the analyzed data.
[0301] A "task ticket" is an electronic ticket that includes task summary, details, deadline, person in charge, and emotion information, which is automatically generated based on analyzed task information.
[0302] A "task management system" is software for registering and managing generated task tickets.
[0303] "Means for notifying" refers to a method or tool for notifying a user of a generated task ticket.
[0304] To implement this invention, it is necessary to build a system using smart glasses and a server equipped with an emotion engine linked to the glasses. This system collects video and audio data from the user's perspective, analyzes it on the server to generate task tickets, performs emotion recognition, registers them in a task management system, and notifies the user.
[0305] Specifically, the system uses the following hardware and software:
[0306] 1. Hardware:
[0307] Smart glasses: Equipped with a built-in camera and microphone.
[0308] Built-in camera: A device for capturing video from the user's point of view in real time.
[0309] Built-in microphone: A device for collecting audio data in real time.
[0310] 2. Software:
[0311] Emotion engine: Software for analyzing the user's emotional state from video and audio data.
[0312] Speech recognition technology: A technology that converts collected voice data into text data.
[0313] Image recognition technology: A technology that extracts task information from video data.
[0314] Natural language processing technology: A technology that analyzes task information from text data.
[0315] Task management system: Software for managing generated task tickets.
[0316] Internet connection: A network environment for sending collected data to a server.
[0317] Here is an example of how the system works:
[0318] Data collection:
[0319] The smart glasses, which function as a device, use a built-in camera to capture images from the user's perspective in real time. The built-in microphone also collects the user's conversation and surrounding sounds. The emotion engine uses this data to analyze the user's emotions. For example, the device simultaneously records the contents of the whiteboard and conversation during a meeting, and captures the task information on the PC screen when the user is looking at it. Furthermore, the device analyzes emotions from the user's facial expressions and tone of voice when they are speaking.
[0320] Data transmission:
[0321] The device sends the collected video data, audio data, and emotion analysis data to a server via the Internet. The data is compressed and encrypted the moment it is collected, improving transmission efficiency and ensuring the security of the information. For example, the device can use a Wi-Fi connection to send video, audio, and emotion analysis data to a server. Similar data transmission is possible even when you are away from home.
[0322] Data Analysis:
[0323] The server inputs the received voice data into a voice recognition engine, which converts the voice into text data. Video data is also input into an image recognition engine, which extracts information related to the task. Natural language processing technology is then used to analyze the task information from this text data. At the same time, an emotion engine analyzes the user's emotions from the voice and video data. For example, if voice data such as "Write a report and submit it by next Monday" is converted into text, the task elements "Write a report," "Submit," and "Next Monday" are extracted. A note written on a whiteboard from the video data saying "Deadline: May 20th" is also analyzed, and the deadline is recognized. The emotion engine analyzes whether the user's voice sounds tense or if they are smiling.
[0324] Task Ticket Generation:
[0325] The server automatically generates a task ticket based on the analyzed task information. The task ticket includes the task summary, details, deadline, assignee, and emotion information, allowing for more context-specific responses. For example, a task ticket containing "Report creation," "Submission deadline: next Monday," "Assignee: User," and "User emotion: Tension" is automatically generated.
[0326] Ticket notifications and management:
[0327] The server registers the generated task ticket in the task management system and notifies the user. The user receives the notification, accesses the task management system to check the generated task ticket, and can make corrections or enter additional information as necessary. It is also possible to adjust the priority of the task and how to respond based on emotional information. For example, the server notifies the user that a "New Task Ticket: Create Report" has been generated, and the user opens the task management app to check the ticket contents, enter notes or supplementary information as necessary, and manage the progress of the task. Taking into account the user's emotional information, the system can respond by lowering the priority of the task if the user is feeling nervous.
[0328] Example prompt sentence:
[0329] "Please analyze the data recorded from the whiteboard contents and conversations during meetings acquired by smart glasses, and extract task information. For example, please capture content that includes information such as 'Deadline: May 20th', and generate a task ticket based on the deadline."
[0330] "Convert task information from user conversation data into text data, and automatically generate task tickets that include emotional data based on a sentiment analysis engine."
[0331] This allows users to efficiently manage a variety of tasks that arise during meetings, conversations, and work, and enables detailed task responses that take into account the user's emotional state. This significantly reduces the burden and errors of manual task entry, improving users' overall work efficiency and satisfaction.
[0332] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0333] Step 1: Data collection
[0334] The device uses a built-in camera to capture video data from the user's perspective in real time. The device's built-in microphone is used to collect surrounding sounds and the user's conversation. The emotion engine receives the collected video and audio data as input and analyzes the user's emotions. For example, it collects text written on a whiteboard and the content of conversations during a meeting, and also obtains emotional information from the user's facial expressions and tone of voice. As a result, collected video data, audio data, and emotional information are generated.
[0335] Step 2: Send data
[0336] The device sends the compressed and encrypted collected data (video data, audio data, and emotional information) to a server via the Internet. In this process, the collected data is sent as input, and the compressed and encrypted data is transferred to the server as output. For example, the data can be sent using a Wi-Fi connection or a mobile network to reach the server safely and efficiently.
[0337] Step 3: Data analysis
[0338] The server inputs the received voice data into a voice recognition engine and converts the user's speech into text data. In parallel, it inputs the video data into an image recognition engine and extracts task-related information from the video. Then, using natural language processing technology, it analyzes the task information from this text data. The emotion engine reanalyzes the user's emotions from the voice and video data. For example, voice data such as "Write and submit a report by next Monday" is converted into text data, and the task elements "Write a report," "Submit," and "Next Monday" are extracted. At the same time, the whiteboard note "Deadline: May 20th" is analyzed and recognized from the video data, and tension is detected from the user's tone of voice. As a result, task information and emotion information are generated.
[0339] Step 4: Task Ticket Generation
[0340] The server automatically generates a task ticket based on the analyzed task information and emotion information. The task ticket includes a summary of the task, details, deadline, person in charge, emotion information, etc. For example, a task ticket containing information such as "Report creation," "Submission deadline: next Monday," "Person in charge: User," and "User emotion: Tension" is created. As a result, a task ticket is generated that reflects the user's emotional state.
[0341] Step 5: Ticket notification and management
[0342] The server registers the generated task ticket in the task management system and notifies the user. The user receives the notification and accesses the task management system to check the generated task ticket. They can modify the task ticket or enter additional information as necessary. It is also possible to adjust the priority and response method of the task based on emotional information. For example, the user is notified that a "New Task Ticket: Create Report" has been created, and the user opens the task management app to check the contents of the ticket and enter notes or supplementary information as necessary. In addition, if the user is feeling nervous, the task priority may be lowered. As a result, efficient task management that takes emotional information into consideration is achieved.
[0343] (Application example 2)
[0344] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0345] Modern factories are becoming increasingly automated to increase productivity, but there are still many situations in which human workers are required. Furthermore, a worker's emotional state and stress level can directly affect work efficiency and quality, but there is still a lack of systems that can detect this in real time and provide work support. Therefore, there is a need for work support and task management that takes into account the emotions of workers.
[0346] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0347] In this invention, the server includes means for capturing video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, means for notifying the user of the generated task tickets, means for analyzing the user's emotional data and adjusting task priorities and response methods based on the emotional data, and means for a robot that receives a task ticket to automatically provide work support or assistance. This enables work support and task management in real time while taking into account the emotional state of the worker.
[0348] A "built-in camera" is a device that captures video from the user's perspective in real time.
[0349] A "built-in microphone" is a device for collecting audio data in real time.
[0350] A "server" is a computer system that analyzes received data and manages various tasks.
[0351] A "task ticket" is a digital ticket that is automatically generated based on analyzed task information and contains a task summary, details, deadline, responsible person, etc.
[0352] "Emotional data" refers to data relating to the emotional state of a user analyzed from their voice and video.
[0353] A "robot" is a mechanical device used to provide support and assistance in work within a factory.
[0354] "Speech recognition technology" is a technology for converting voice data into text data.
[0355] "Image recognition technology" is a technology for extracting useful information from video data.
[0356] "Natural language processing technology" is a technology for analyzing intent and content from text data.
[0357] A "task management system" is a system that manages generated task tickets and allows users to access and modify them.
[0358] "Work support" refers to an assistive action that the robot performs based on a task ticket.
[0359] To implement this invention, the following components are required: smart glasses worn by workers in the factory, a server that performs data analysis, and a robot that supports the work.
[0360] Hardware and Software Configuration
[0361] Smart Glasses
[0362] Smart glasses, such as Google Glass Enterprise Edition, are equipped with built-in cameras and microphones to collect real-time video and audio data from the worker's perspective.
[0363] server
[0364] The server is a high-performance data analysis server, and can use Amazon Web Services' EC2, etc. The server uses the following software:
[0365] Speech recognition engine: Uses Google Cloud Speech-to-Text to convert voice data into text data.
[0366] Image recognition engine: Extracts task information from video data using OpenCV and TensorFlow.
[0367] Natural language processing technology: BERT is used to analyze task information from text data.
[0368] Emotion recognition engine: Affectiva is used to analyze emotional information from audio and video data.
[0369] Task management system: Use JIRA or similar to manage generated task tickets.
[0370] robot
[0371] The Universal Robots UR series is suitable as a work support robot for use in factories.
[0372] Examples of data collection
[0373] The smart glasses, which serve as the device, capture video from the worker's point of view in real time and collect audio data with a built-in microphone. For example, they can record video of a worker's hands as they assemble parts, along with audio of a worker saying, "This part doesn't fit properly."
[0374] Specific examples of data transmission
[0375] The device sends the collected video and audio data to a server via a Wi-Fi connection, where the data is compressed, encrypted, and securely transferred.
[0376] Specific examples of data analysis
[0377] The server converts the received voice data into text data using Google Cloud Speech-to-Text, and analyzes the speech, "The left screw won't tighten properly." Next, the video data is analyzed using OpenCV and TensorFlow to extract the screw installation position and the condition of the parts. At the same time, the emotion recognition engine Affectiva detects stress from the worker's voice.
[0378] Example of task ticket generation
[0379] The server automatically generates a task ticket based on the analyzed task information and emotion information. For example, a task ticket such as "Screw tightening", "Duties: Operations team", "Deadline: Immediately", and "Emotion: Stress" may be generated.
[0380] Task ticket notification and management example
[0381] The generated task ticket is registered in the task management system JIRA and notified to the work support robot. Based on this information, the robot displays a step-by-step video guide to the worker and supports the procedure. If the worker is feeling stressed, the robot automatically provides support for the work.
[0382] Examples of prompt statements
[0383] Invention details: A system that analyzes data collected by smart glasses on a server and generates task tickets.
[0384] Usage scenario: A robot that supports work in a factory operates based on video, audio, and emotional data.
[0385] Input data example: Video, audio, and emotion data sent from smart glasses.
[0386] Example of output task ticket: "Screw tightening" "Worker: Operations team" "Deadline: Immediately" "Emotion: Stress"
[0387] Thus, the present invention provides a system that allows for real-time work support and task management while taking into account the emotional state of the worker.
[0388] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0389] Step 1:
[0390] The device (smart glasses) collects video and audio data.
[0391] Input: Video from the user's perspective and surrounding audio data
[0392] Output: Real-time collected video and audio data
[0393] How it works: The smart glasses' built-in camera captures video from the user's point of view, and the built-in microphone collects audio, recording, for example, a worker assembling parts and their conversations.
[0394] Step 2:
[0395] The device sends the data to the server
[0396] Input: Collected video and audio data
[0397] Output: Video and audio data transferred to the server
[0398] How it works: The smart glasses send the data they collect to a server via Wi-Fi. The data is compressed and encrypted for secure transmission. For example, a video of the worker's hands and a voice message saying, "This part isn't fitting properly" are sent.
[0399] Step 3:
[0400] The server converts the audio data into text data.
[0401] Input: Audio data sent from the device
[0402] Output: Character data
[0403] Specific operation: The server converts voice data into text data using Google Cloud Speech-to-Text. For example, a speech saying "I can't tighten the left screw properly" is converted into text data.
[0404] Step 4:
[0405] The server analyzes the video data and extracts task information.
[0406] Input: Video data sent from the device
[0407] Output: Task information
[0408] How it works: The server uses OpenCV and TensorFlow to analyze the video data and extract information relevant to the task, such as the position of screws or the condition of parts.
[0409] Step 5:
[0410] The server analyzes the emotional data
[0411] Input: Audio and video data
[0412] Output: Emotion data
[0413] How it works: The server uses Affectiva to analyze audio and video data and determine the user's emotions. For example, it can detect that a worker is feeling stressed.
[0414] Step 6:
[0415] The server analyzes the task information and generates a task ticket.
[0416] Input: Text data, task information extracted from video, emotion data
[0417] Output: Task ticket
[0418] Specific operation: The server analyzes the task information using BERT and generates a task ticket based on that information. For example, a task ticket with the following content might be created: "Screw tightening," "Duty: Operations team," "Deadline: Immediate," and "Emotion: Stress."
[0419] Step 7:
[0420] The server registers the task ticket in the task management system and notifies it.
[0421] Input: Generated task ticket
[0422] Output: Task tickets registered in the task management system
[0423] Specific operation: The task ticket generated by the server is registered in the task management system (e.g. JIRA) and this information is then notified to the robot.
[0424] Step 8:
[0425] Robots provide support and assistance for tasks
[0426] Input: Task ticket notified from the server
[0427] Output: Assistance performed
[0428] Specific behavior: The robot performs work support and assistance based on the contents of the task ticket, for example, displaying a step-by-step video guide when a worker is feeling stressed or bringing parts for the worker.
[0429] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0430] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0431] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0432] [Second embodiment]
[0433] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0434] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0435] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0436] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0437] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0438] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0439] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0440] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0441] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0442] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0443] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0444] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0445] To implement this invention, we will build a system using smart glasses and a server connected to them. This system collects video and audio data from the user's point of view, analyzes it on the server, generates task tickets, and registers them in a task management system.
[0446] Program processing overview
[0447] Data collection
[0448] The device uses the smart glasses' built-in camera to capture real-time video from the user's perspective, and a built-in microphone to collect the user's conversation and surrounding sounds, allowing for accurate recording of task information during meetings and conversations.
[0449] Example: The device simultaneously records whiteboard content and conversations during meetings. When the user is looking at a PC screen, the device also captures on-screen task information.
[0450] Data transmission
[0451] The devices transmit the collected video and audio data in real time to a server over an internet connection, where the data is compressed for efficient transmission and encrypted to ensure the information is secure.
[0452] Example: A device uses a Wi-Fi connection to send video and audio data to a server. Data can be sent even when you are away from home.
[0453] Data analysis
[0454] The server inputs the received video data into an image recognition engine to extract important elements including task information. At the same time, it inputs the audio data into a voice recognition engine and converts it into text data. Then, using natural language processing technology, it analyzes the task information from this text data.
[0455] Example: The server converts audio data such as "Write a report and submit it by next Monday" into text and extracts the task elements "Write a report," "Submit," and "Next Monday." It also analyzes a whiteboard note from video data that reads "Deadline: May 20th" and recognizes the deadline.
[0456] Generate a Task Ticket
[0457] The server automatically generates a task ticket based on the parsed task information, which includes information such as the task summary, details, deadline, and assignee, allowing users to manage tasks accurately and quickly without any manual intervention.
[0458] Example: A task ticket with the following name will be automatically generated: "Create a report", "Submission deadline: next Monday", "Responsible person: User".
[0459] Ticket notifications and management
[0460] The server registers the generated task ticket in the task management system and has a notification means for notifying the user. The user receives the notification, accesses the task management system, checks the generated task ticket, and can correct or enter additional information as necessary.
[0461] Example: The server notifies the user that a "New Task Ticket: Create Report" has been created, and the user opens a task management app to check the contents of the ticket. If necessary, they can enter notes or supplementary information and manage the progress of the task.
[0462] This system allows users to efficiently manage various tasks that arise during meetings, conversations, and work, significantly reducing the burden and errors of manual ticket creation.
[0463] The processing flow will be explained below.
[0464] Step 1:
[0465] The device uses a built-in camera to capture video from the user's point of view in real time, including the contents of the whiteboard, presentation slides, and PC screen during the meeting.
[0466] Step 2:
[0467] The device uses a built-in microphone to collect audio data in real time, recording the user's conversations and the sounds of the surrounding area.
[0468] Step 3:
[0469] The devices transmit the collected video and audio data to a server over the Internet, where the data is pre-compressed and encrypted to ensure efficient and secure transmission.
[0470] Step 4:
[0471] The server inputs the received voice data into a speech recognition engine and converts the voice into text data, making the conversation available in text format.
[0472] Step 5:
[0473] The server inputs the received video data into an image recognition engine and extracts task-related information from the video, for example, identifying text written on a whiteboard or task information displayed on a PC screen.
[0474] Step 6:
[0475] The server uses natural language processing technology to analyze the converted text data and identify task information, such as the task content, deadline, and person in charge, from the conversation and video data.
[0476] Step 7:
[0477] The server automatically generates a task ticket based on the analyzed task information. The generated task ticket includes information such as the task summary, details, deadline, and person in charge.
[0478] Step 8:
[0479] The server registers the generated task ticket in the task management system, which allows tasks to be centrally managed and users to check them later.
[0480] Step 9:
[0481] The server will notify the user of the newly created task ticket, via email, push notification, or other methods.
[0482] Step 10:
[0483] The user receives a notification and accesses the task management system to view the generated task ticket, enter additional notes and details as needed, and manage the task's progress.
[0484] Example 1
[0485] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0486] Conventional task management systems require task information to be entered manually, a time-consuming process that poses a high risk of input errors. Furthermore, there was no way to efficiently manage tasks that arise during meetings or conversations, which meant there was a risk of important information being overlooked or tasks being left out.
[0487] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0488] In this invention, the server includes means for capturing video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for compressing and encrypting the transmitted data, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, and means for notifying the user of the generated task tickets. This makes it possible to reliably collect and analyze task information in real time even during meetings and conversations, and automatically generate and manage task tickets.
[0489] An "integrated camera" is a camera that is built into a device and has the ability to capture video from the user's perspective in real time.
[0490] An "internal microphone" is a microphone that is built into a device and has the ability to collect audio data in real time.
[0491] A "server" is a computer system on a network that receives collected data and performs analysis and task ticket generation.
[0492] "Real-time" is a time processing concept that means data is generated, transmitted, and analyzed almost immediately.
[0493] "Speech recognition technology" is a technology that analyzes voice data and converts the content into text data.
[0494] "Character data" is text information obtained from voice data using voice recognition technology.
[0495] "Image recognition technology" is a technology that extracts specific information and patterns from video data.
[0496] "Natural language processing technology" is a general term for algorithms and technologies that analyze text data and understand, classify, and extract its content.
[0497] A "task ticket" is a manageable unit of a task that includes information such as task details, deadlines, and people in charge.
[0498] "Task Management System" means software or platform that registers and tracks generated task tickets and assists users in managing their tasks.
[0499] "Compression" is the process of reducing the size of data, a technique used to improve transmission efficiency.
[0500] "Encryption" is a technology that converts data to protect it and prevent unauthorized access or tampering during transmission.
[0501] A "push notification" is a notification message sent in real time from a server to a user's device.
[0502] To implement this invention, it is necessary to build a system using smart glasses and a server connected to them. The present invention collects video and audio data from the user's perspective, analyzes it on the server, generates task tickets, and registers them in a task management system.
[0503] First, the user puts on the smart glasses. The smart glasses are equipped with a built-in camera and microphone, which are used to collect video and audio data in real time. Specifically, the camera captures the video the user is looking at, and the microphone collects the conversations of the user and those around them. Typical smart glasses products can be used as the specific hardware for this purpose.
[0504] The device then transmits the collected video and audio data in real time to a server. During this transmission process, the data is compressed and encrypted. An internet connection is used as the communication medium. At this stage, data security and efficient transfer are crucial.
[0505] The server inputs the received video data into an image recognition engine and extracts key elements, including task information. Software used can include Google Cloud Vision API for image recognition, Google Cloud Speech-to-Text for voice recognition, and NLP technologies such as spaCy and Transformers for natural language processing. For example, if voice data such as "Write a report and submit it by next Monday" is sent to the server, it is converted into text using voice recognition technology, and task elements such as "Write a report," "Submit," and "Next Monday" are extracted. Furthermore, a whiteboard note from the video data stating "Deadline: May 20th" is analyzed, and the deadline is recognized.
[0506] Next, the server automatically generates a task ticket based on the analyzed task information. This task ticket includes the task summary, details, deadline, person in charge, etc. For example, specific information such as "Create a report," "Submission deadline: next Monday," and "Person in charge: User" is written in the task ticket.
[0507] The generated task ticket is registered in the task management system by the server. Commonly used platforms such as Trello and Jira can be used as the task management system. The server notifies the user of the generated task ticket. Possible notification methods include push notification and email.
[0508] Finally, the user accesses the task management system based on the received notification and checks the generated task ticket. The user can enter corrections or additional information into the task ticket as necessary to manage the progress of the task.
[0509] Specific examples
[0510] For example, if a user is told during a meeting to "write a report and submit it by next Monday," the user's video and audio will be collected through smart glasses. The device will then send the data to a server, which will then use speech recognition and natural language processing to extract task information. This will automatically generate a task ticket in the task management system with the following information: "Create a report," "Submission deadline: next Monday," and "Responsible party: User."
[0511] Prompt Sentence Examples
[0512] "Interpret the notes on the whiteboard during the meeting and extract important task information. For example, generate a task ticket with the content 'Submit report due May 20th'."
[0513] This system makes it possible to efficiently manage the various tasks that arise during meetings and conversations, significantly reducing the effort required to manually enter tasks and the risk of entering them incorrectly.
[0514] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0515] Step 1:
[0516] The device uses the camera built into the smart glasses to capture video from the user's point of view in real time, while simultaneously using the built-in microphone to collect the user's conversation and surrounding sounds. The input data is video data and audio data, and the output data is the collected raw data.
[0517] Specific behavior:
[0518] The device activates the smart glasses camera.
[0519] The device begins collecting audio data using the built-in microphone.
[0520] When a user is looking at notes on a whiteboard during a meeting, the terminal records the contents of the whiteboard as video.
[0521] When a user speaks, the microphone in the smart glasses picks up the sound.
[0522] Step 2:
[0523] The terminal transmits collected video and audio data to the server in real time. The data is compressed for efficiency and encrypted for security. The input data is the collected raw data. The output data is the compressed and encrypted data.
[0524] Specific behavior:
[0525] The device compresses the video and audio data.
[0526] The device encrypts the data.
[0527] The device sends the data to the server through an internet connection.
[0528] The server decompresses and decrypts the received data.
[0529] Step 3:
[0530] The server inputs the received video data into an image recognition engine to extract important elements including task information. At the same time, it inputs the audio data into a speech recognition engine and converts it into text data. The input data is raw data that has been decompressed and decrypted. The output data is the analyzed task information.
[0531] Specific behavior:
[0532] The server sends the video data to an image recognition engine, for example, using the Google Cloud Vision API.
[0533] The server sends the voice data to a speech recognition engine and converts it into text data. For example, Google Cloud Speech-to-Text is used.
[0534] The server sends the converted text data to a natural language processing engine, which analyzes and extracts task information. For example, we use spaCy.
[0535] The server extracts task details such as "write a report," "submit," and "next Monday" from the voice data.
[0536] The server analyzes the whiteboard note "Deadline: May 20th" from the video data and recognizes the task deadline.
[0537] Step 4:
[0538] The server automatically generates a task ticket based on the parsed task information. The task ticket contains information such as the task summary, details, deadline, and person in charge. The input data is the parsed task information. The output data is the generated task ticket.
[0539] Specific behavior:
[0540] The server automatically generates a task ticket based on the analyzed task information.
[0541] The server writes information such as "Create report," "Submission deadline: next Monday," and "Responsible person: User" into the task ticket.
[0542] Step 5:
[0543] The server registers the generated task ticket in the task management system and notifies the user. The input data is the generated task ticket. The output data is the notified task ticket information.
[0544] Specific behavior:
[0545] The server registers the task ticket in a task management system, such as Trello or Jira.
[0546] The server notifies the user of the new task ticket creation via push notification or email.
[0547] A user accesses the task management system to view the new task ticket and enter notes or additional information as needed.
[0548] (Application example 1)
[0549] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0550] In modern factory environments, many tasks are still performed manually, and efficiency and accuracy are essential. However, workers spend a lot of time manually creating and managing task tickets, resulting in reduced productivity. Furthermore, there is no system in place to accurately record the various tasks that occur during work and instantly notify robots. This disrupts the smooth progress of work within the factory, making it difficult to address any errors or delays that arise.
[0551] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0552] In this invention, the server includes means for capturing video from a user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, means for notifying the generated task tickets to users, means for having workers wear smart glasses in a factory environment to collect task details, send the collected data to the server for analysis, and initiate task notification to factory robots, and means for supporting workers to receive work instructions in real time and perform tasks efficiently, thereby enabling automation and improved accuracy of task management within the factory.
[0553] "Built-in camera" refers to a camera module built into an apparatus or device that captures video from the user's perspective in real time.
[0554] "Built-in microphone" refers to a microphone built into an apparatus or device for collecting audio data in real time.
[0555] A "server" refers to a computer system that processes data over a network and has storage and analysis capabilities.
[0556] "Speech recognition technology" refers to technology that analyzes voice data and converts that voice into text data.
[0557] "Image recognition technology" refers to technology for analyzing video data and extracting specific information from it.
[0558] "Natural language processing technology" refers to technology that analyzes text data and extracts and understands information based on human language.
[0559] A "task ticket" is an electronic card or format that compiles information about a specific task, including the task name, details, deadline, and assignee.
[0560] A "task management system" refers to software for registering and managing generated task tickets.
[0561] "Smart glasses" refers to a wearable eyeglass-type device that has augmented reality and information display functions.
[0562] "Factory robot" refers to a robot used to automate production tasks in a factory.
[0563] "Worker" refers to a person who actually performs manual work or supervises in a factory or production site.
[0564] To implement this invention, a system using a server linked to smart glasses is constructed as follows.
[0565] Data collection
[0566] Device:
[0567] The user wears the smart glasses, and while working, the built-in camera captures video from the user's point of view in real time. The built-in microphone collects audio data. Examples of hardware used to collect this data include the built-in camera from Logitech and the microphone from Blue Yeti.
[0568] Specific examples
[0569] Workers on the factory assembly line wear smart glasses and perform their tasks while receiving verbal instructions on part installation procedures and important points to note, which then record the work procedures as visual and audio data.
[0570] Data transmission
[0571] Device:
[0572] The collected video and audio data is sent to a server via the Internet, where it is compressed and encrypted to ensure efficient and secure transmission.
[0573] Data analysis
[0574] server:
[0575] The server inputs the received video data into an image recognition engine (e.g., OpenCV, TensorFlow) to extract key elements including task information. It also inputs the audio data into a speech recognition engine (e.g., Google Speech-to-Text API) to convert it into text data. It then uses natural language processing technology (e.g., NLTK, spaCy) to analyze the text data for specific task information.
[0576] Specific examples
[0577] The server converts the voice data, "Tighten this bolt according to step 7," into text and extracts the task elements "bolt tightening task" and "step 7." It also analyzes and recognizes signs from the video data that include the instruction "inspect after task completion."
[0578] Task ticket generation and notification
[0579] server:
[0580] A task ticket is automatically generated based on the analyzed task information. The task ticket contains information such as the task summary, details, deadline, and person in charge. The generated task ticket is registered in the task management system, and its contents are notified to the user.
[0581] Specific examples
[0582] Task tickets for "Bolt tightening," "Step 7," and "Inspect after completion" are automatically generated and displayed on the worker's smart glasses. The same information is also sent to the factory robot, which automatically supports the necessary tasks.
[0583] Real-time support for work instructions
[0584] Terminals and robots:
[0585] The smart glasses help workers receive real-time work instructions to carry out their work efficiently. They also display the next steps and important points to note, and send instructions to factory robots.
[0586] Prompt Sentence Examples
[0587] If you're having a discussion in a conference room:
[0588] I need to write and submit the report by next Monday.
[0589] For factory assembly lines:
[0590] "When you assemble this part to the next block, you need to tighten the nuts."
[0591] These processes significantly improve task management and work efficiency within the factory, reducing the burden and errors of manual ticket creation.
[0592] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0593] Program processing flow
[0594] Step 1: Data collection
[0595] Device:
[0596] Using the built-in camera, video data is captured in real time from the user's perspective.
[0597] The built-in microphone is used to collect the user's voice data in real time.
[0598] Input: Video data, audio data
[0599] Output: Collected video and audio files
[0600] Step 2: Send data
[0601] Device:
[0602] The collected video and audio data is sent to a server via the Internet.
[0603] Data is compressed and encrypted for secure and efficient transfer.
[0604] Input: Collected video and audio files
[0605] Output: Data sent to the server
[0606] Step 3: Data analysis (audio)
[0607] server:
[0608] The voice data is input into a voice recognition engine, and the voice is converted into text data.
[0609] Use voice recognition technology (e.g., Google Speech-to-Text API).
[0610] Input: Audio data
[0611] Output: Text data
[0612] Step 4: Data analysis (video)
[0613] server:
[0614] The video data is input into an image recognition engine, and task information is extracted from the video.
[0615] Uses image recognition technology (e.g., OpenCV, TensorFlow).
[0616] Input: Video data
[0617] Output: Extracted task information
[0618] Step 5: Parse task information
[0619] server:
[0620] Text data is input into natural language processing technology to analyze task information.
[0621] Use natural language processing technology (e.g., NLTK, spaCy).
[0622] Input: Text data
[0623] Output: Parsed task information
[0624] Step 6: Generate a task ticket
[0625] server:
[0626] Based on the parsed task information, a task ticket is automatically generated.
[0627] The task information includes a summary of the task, details, deadlines, and the person in charge.
[0628] Input: Parsed task information
[0629] Output: Generated task ticket
[0630] Step 7: Task ticket registration and notification
[0631] server:
[0632] The generated task ticket is registered in the task management system and notified to the user.
[0633] Input: Generated task ticket
[0634] Output: Notify user and register with task management system
[0635] Step 8: Real-time support for work instructions
[0636] Terminals and robots:
[0637] Workers receive work instructions in real time and are supported to carry out their work efficiently.
[0638] The smart glasses display the next steps and precautions to take, and send instructions to factory robots.
[0639] Input: Task ticket information
[0640] Output: Work instructions for smart glasses and factory robots
[0641] Specific operation explanation
[0642] Step 1: The device uses the smart glasses' built-in camera and microphone to collect real-time video and audio from the worker's perspective. The worker explains how to install the parts as they go along, collecting the necessary data.
[0643] Step 2: The collected video and audio data is sent to a server via the internet. The data is compressed and encrypted for efficient and secure transmission.
[0644] Step 3: The server uses a speech recognition engine to convert the voice data into text data. For example, the voice saying "Tighten this bolt according to step 7" is converted into text data.
[0645] Step 4: The server uses a video recognition engine to extract task information from the video data. For example, it can read "Step 7" and "Check after inspection" from signs and labels in the video.
[0646] Step 5: The server uses natural language processing technology to analyze specific task information from the converted text data. For example, it extracts task elements such as "bolt tightening," "step 7," and "inspection after completion of work."
[0647] Step 6: The server automatically generates a task ticket based on the parsed task information. The task ticket includes the task name, details, deadline, and person in charge.
[0648] Step 7: The generated task ticket is registered in the task management system and simultaneously notified to the worker's smart glasses and factory robots. The worker can then view the new task through the smart glasses.
[0649] Step 8: Using smart glasses, workers receive real-time instructions on the next steps and important points to note, allowing them to work efficiently. Factory robots also receive instructions and automatically assist with the designated tasks.
[0650] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0651] To implement this invention, we will build a system using smart glasses and a server equipped with an emotion engine linked to the glasses. This system collects video and audio data from the user's perspective, analyzes it on the server, generates task tickets, performs emotion recognition, registers them in a task management system, and notifies the user.
[0652] Program processing overview
[0653] Data collection
[0654] The device uses the smart glasses' built-in camera to capture real-time video from the user's point of view, and a built-in microphone to collect the user's conversation and surrounding sounds. The emotion engine uses this data to analyze the user's emotions.
[0655] Example: The device simultaneously records the content of the whiteboard and conversations during a meeting. When the user is looking at the PC screen, the device also captures the task information on the screen. Furthermore, the device analyzes the user's emotions from facial expressions and tone of voice while they are speaking.
[0656] Data transmission
[0657] The device transmits the collected video, audio, and emotion analysis data to a server via the internet. This transmission uses an internet connection. The data is compressed to improve transmission efficiency and encrypted to ensure the security of the information.
[0658] Example: A device uses a Wi-Fi connection to send video and audio data, as well as emotion analysis data, to a server. Data can be sent even when the user is away from home.
[0659] Data analysis
[0660] The server inputs the received voice data into a voice recognition engine, which converts the voice into text data. It then inputs the video data into an image recognition engine, which extracts task-related information from the video. It then uses natural language processing technology to analyze the task information from the text data. The emotion engine also analyzes the user's emotions from the voice and video data.
[0661] Example: The server converts voice data such as "Write a report and submit it by next Monday" into text and extracts the task elements "Write a report," "Submit," and "Next Monday." It also analyzes a whiteboard note from video data that reads "Deadline: May 20th" to recognize the deadline. At the same time, the emotion engine analyzes whether the user's voice sounds tense or whether they are smiling.
[0662] Generate a Task Ticket
[0663] The server automatically generates a task ticket based on the parsed task information, which includes the task summary, details, deadline, assignee, and user sentiment information, allowing for more context-sensitive responses.
[0664] Example: A task ticket containing "Write a report", "Submission deadline: next Monday", "Responsible party: User", and "User emotion: Tension" is automatically generated.
[0665] Ticket notifications and management
[0666] The server registers the generated task ticket in the task management system and has a notification means to notify the user. The user receives the notification, accesses the task management system to check the generated task ticket, and can correct or add additional information as necessary. It is also possible to adjust the priority and response method of the task based on the emotional information.
[0667] Example: The server notifies the user that a "New Task Ticket: Create Report" has been created, and the user opens a task management app to check the contents of the ticket. If necessary, the user can enter notes or supplementary information and manage the progress of the task. Taking into account the user's emotional information, if the user is feeling nervous, the system will take appropriate measures such as lowering the priority of the task.
[0668] This system allows users to efficiently manage various tasks that arise during meetings, conversations, and work, and enables detailed task responses that take into account the user's emotional state, significantly reducing the burden and errors of manual task entry and improving users' overall work efficiency and satisfaction.
[0669] The processing flow will be explained below.
[0670] Step 1:
[0671] The device uses a built-in camera to capture real-time video from the user's point of view, including the contents of the whiteboard, presentation slides, and PC screen during the meeting.
[0672] Step 2:
[0673] The device uses a built-in microphone to collect voice data in real time, recording the user's conversation and surrounding sounds, and an emotion engine analyzes the tone and tempo of the voice to recognize the user's emotions.
[0674] Step 3:
[0675] The device then sends the collected video, audio, and emotion analysis data to a server via the internet. The data is compressed to improve transmission efficiency and encrypted to ensure the security of the information.
[0676] Step 4:
[0677] The server inputs the received voice data into a speech recognition engine and converts the voice into text data, making the conversation available in text format.
[0678] Step 5:
[0679] The server inputs the received video data into an image recognition engine and extracts task-related information from the video, for example, identifying text written on a whiteboard or task information displayed on a PC screen.
[0680] Step 6:
[0681] The server uses natural language processing technology to analyze task information from the converted text data. Specifically, it identifies the task content, deadline, and person in charge from the conversation and video data. In addition, an emotion engine analyzes the user's emotions from the audio and video data.
[0682] Step 7:
[0683] The server automatically generates a task ticket based on the analyzed task information and the user's emotion information. The generated task ticket includes the task summary, details, deadline, assignee, and user's emotion information.
[0684] Step 8:
[0685] The server registers the generated task ticket in the task management system. At that time, it is also possible to set the priority of the task based on the user's emotional information.
[0686] Step 9:
[0687] The server will notify the user of the newly created task ticket, via email, push notification, or other methods.
[0688] Step 10:
[0689] Users receive a notification and access their task management system to view the generated task ticket, enter additional notes and details as needed, manage the task's progress, and even adjust the task's priority based on sentiment information.
[0690] Example 2
[0691] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0692] In today's business environment, where a variety of tasks are being carried out simultaneously, manually inputting and managing tasks by users is laborious and prone to errors. Furthermore, task management that takes into account the user's emotional state is difficult, resulting in reduced work efficiency and satisfaction. Current task management systems are unable to automatically generate tasks in real time or manage tasks that take into account emotional information, so improvements are needed to address this issue.
[0693] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for compressing and encrypting the collected video data and audio data and transmitting them to the server, means for converting audio data into text data using voice recognition technology, means for extracting task information from video data using image recognition technology, means for analyzing task information from text data using natural language processing technology, means for analyzing user emotions from the collected video data and audio data using emotion analysis technology, means for automatically generating task tickets based on task information, means for adding emotion information to the generated task tickets, means for registering the generated task tickets in a task management system, and means for notifying the user of the generated task tickets. This enables efficient and accurate task management that also takes the user's emotional state into consideration.
[0694] A "built-in camera" is a photographing device built into a terminal to capture images from the user's point of view in real time.
[0695] A "built-in microphone" is a voice collection device built into a terminal for collecting voice data in real time.
[0696] "Collected Video Data" means visual information captured in real time by the built-in camera.
[0697] "Collected Audio Data" means audio information collected in real time by the built-in microphone.
[0698] "Means of compressing, encrypting and transmitting" refers to the technical means of compressing collected data for efficient transmission, encrypting it to ensure the security of communications and transmitting it to a server via the Internet.
[0699] "Speech recognition technology" is a technology that analyzes collected voice data and converts it into text data.
[0700] "Character data" is voice information analyzed by voice recognition technology expressed in text format.
[0701] "Image recognition technology" is a technology that analyzes collected video data and extracts task information from the video.
[0702] "Natural language processing technology" is a technology that analyzes task information from text data and understands meaning and intent.
[0703] "Emotion analysis technology" is a technology that analyzes collected video and audio data to recognize the user's emotional state.
[0704] "Task information" refers to information related to tasks and instructions extracted from the analyzed data.
[0705] A "task ticket" is an electronic ticket that includes task summary, details, deadline, person in charge, and emotion information, which is automatically generated based on analyzed task information.
[0706] A "task management system" is software for registering and managing generated task tickets.
[0707] "Means for notifying" refers to a method or tool for notifying a user of a generated task ticket.
[0708] To implement this invention, it is necessary to build a system using smart glasses and a server equipped with an emotion engine linked to the glasses. This system collects video and audio data from the user's perspective, analyzes it on the server to generate task tickets, performs emotion recognition, registers them in a task management system, and notifies the user.
[0709] Specifically, the system uses the following hardware and software:
[0710] 1. Hardware:
[0711] Smart glasses: Equipped with a built-in camera and microphone.
[0712] Built-in camera: A device for capturing video from the user's point of view in real time.
[0713] Built-in microphone: A device for collecting audio data in real time.
[0714] 2. Software:
[0715] Emotion engine: Software for analyzing the user's emotional state from video and audio data.
[0716] Speech recognition technology: A technology that converts collected voice data into text data.
[0717] Image recognition technology: A technology that extracts task information from video data.
[0718] Natural language processing technology: A technology that analyzes task information from text data.
[0719] Task management system: Software for managing generated task tickets.
[0720] Internet connection: A network environment for sending collected data to a server.
[0721] Here is an example of how the system works:
[0722] Data collection:
[0723] The smart glasses, which function as a device, use a built-in camera to capture images from the user's perspective in real time. The built-in microphone also collects the user's conversation and surrounding sounds. The emotion engine uses this data to analyze the user's emotions. For example, the device simultaneously records the contents of the whiteboard and conversation during a meeting, and captures the task information on the PC screen when the user is looking at it. Furthermore, the device analyzes emotions from the user's facial expressions and tone of voice when they are speaking.
[0724] Data transmission:
[0725] The device sends the collected video data, audio data, and emotion analysis data to a server via the Internet. The data is compressed and encrypted the moment it is collected, improving transmission efficiency and ensuring the security of the information. For example, the device can use a Wi-Fi connection to send video, audio, and emotion analysis data to a server. Similar data transmission is possible even when you are away from home.
[0726] Data Analysis:
[0727] The server inputs the received voice data into a voice recognition engine, which converts the voice into text data. Video data is also input into an image recognition engine, which extracts information related to the task. Natural language processing technology is then used to analyze the task information from this text data. At the same time, an emotion engine analyzes the user's emotions from the voice and video data. For example, if voice data such as "Write a report and submit it by next Monday" is converted into text, the task elements "Write a report," "Submit," and "Next Monday" are extracted. A note written on a whiteboard from the video data saying "Deadline: May 20th" is also analyzed, and the deadline is recognized. The emotion engine analyzes whether the user's voice sounds tense or if they are smiling.
[0728] Task Ticket Generation:
[0729] The server automatically generates a task ticket based on the analyzed task information. The task ticket includes the task summary, details, deadline, assignee, and emotion information, allowing for more context-specific responses. For example, a task ticket containing "Report creation," "Submission deadline: next Monday," "Assignee: User," and "User emotion: Tension" is automatically generated.
[0730] Ticket notifications and management:
[0731] The server registers the generated task ticket in the task management system and notifies the user. The user receives the notification, accesses the task management system to check the generated task ticket, and can make corrections or enter additional information as necessary. It is also possible to adjust the priority of the task and how to respond based on emotional information. For example, the server notifies the user that a "New Task Ticket: Create Report" has been generated, and the user opens the task management app to check the ticket contents, enter notes or supplementary information as necessary, and manage the progress of the task. Taking into account the user's emotional information, the system can respond by lowering the priority of the task if the user is feeling nervous.
[0732] Example prompt sentence:
[0733] "Please analyze the data recorded from the whiteboard contents and conversations during meetings acquired by smart glasses, and extract task information. For example, please capture content that includes information such as 'Deadline: May 20th', and generate a task ticket based on the deadline."
[0734] "Convert task information from user conversation data into text data, and automatically generate task tickets that include emotional data based on a sentiment analysis engine."
[0735] This allows users to efficiently manage a variety of tasks that arise during meetings, conversations, and work, and enables detailed task responses that take into account the user's emotional state. This significantly reduces the burden and errors of manual task entry, improving users' overall work efficiency and satisfaction.
[0736] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0737] Step 1: Data collection
[0738] The device uses a built-in camera to capture video data from the user's perspective in real time. The device's built-in microphone is used to collect surrounding sounds and the user's conversation. The emotion engine receives the collected video and audio data as input and analyzes the user's emotions. For example, it collects text written on a whiteboard and the content of conversations during a meeting, and also obtains emotional information from the user's facial expressions and tone of voice. As a result, collected video data, audio data, and emotional information are generated.
[0739] Step 2: Send data
[0740] The device sends the compressed and encrypted collected data (video data, audio data, and emotional information) to a server via the Internet. In this process, the collected data is sent as input, and the compressed and encrypted data is transferred to the server as output. For example, the data can be sent using a Wi-Fi connection or a mobile network to reach the server safely and efficiently.
[0741] Step 3: Data analysis
[0742] The server inputs the received voice data into a voice recognition engine and converts the user's speech into text data. In parallel, it inputs the video data into an image recognition engine and extracts task-related information from the video. Then, using natural language processing technology, it analyzes the task information from this text data. The emotion engine reanalyzes the user's emotions from the voice and video data. For example, voice data such as "Write and submit a report by next Monday" is converted into text data, and the task elements "Write a report," "Submit," and "Next Monday" are extracted. At the same time, the whiteboard note "Deadline: May 20th" is analyzed and recognized from the video data, and tension is detected from the user's tone of voice. As a result, task information and emotion information are generated.
[0743] Step 4: Task Ticket Generation
[0744] The server automatically generates a task ticket based on the analyzed task information and emotion information. The task ticket includes a summary of the task, details, deadline, person in charge, emotion information, etc. For example, a task ticket containing information such as "Report creation," "Submission deadline: next Monday," "Person in charge: User," and "User emotion: Tension" is created. As a result, a task ticket is generated that reflects the user's emotional state.
[0745] Step 5: Ticket notification and management
[0746] The server registers the generated task ticket in the task management system and notifies the user. The user receives the notification and accesses the task management system to check the generated task ticket. They can modify the task ticket or enter additional information as necessary. It is also possible to adjust the priority and response method of the task based on emotional information. For example, the user is notified that a "New Task Ticket: Create Report" has been created, and the user opens the task management app to check the contents of the ticket and enter notes or supplementary information as necessary. In addition, if the user is feeling nervous, the task priority may be lowered. As a result, efficient task management that takes emotional information into consideration is achieved.
[0747] (Application example 2)
[0748] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0749] Modern factories are becoming increasingly automated to increase productivity, but there are still many situations in which human workers are required. Furthermore, a worker's emotional state and stress level can directly affect work efficiency and quality, but there is still a lack of systems that can detect this in real time and provide work support. Therefore, there is a need for work support and task management that takes into account the emotions of workers.
[0750] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0751] In this invention, the server includes means for capturing video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, means for notifying the user of the generated task tickets, means for analyzing the user's emotional data and adjusting task priorities and response methods based on the emotional data, and means for a robot that receives a task ticket to automatically provide work support or assistance. This enables work support and task management in real time while taking into account the emotional state of the worker.
[0752] A "built-in camera" is a device that captures video from the user's perspective in real time.
[0753] A "built-in microphone" is a device for collecting audio data in real time.
[0754] A "server" is a computer system that analyzes received data and manages various tasks.
[0755] A "task ticket" is a digital ticket that is automatically generated based on analyzed task information and contains a task summary, details, deadline, responsible person, etc.
[0756] "Emotional data" refers to data relating to the emotional state of a user analyzed from their voice and video.
[0757] A "robot" is a mechanical device used to provide support and assistance in work within a factory.
[0758] "Speech recognition technology" is a technology for converting voice data into text data.
[0759] "Image recognition technology" is a technology for extracting useful information from video data.
[0760] "Natural language processing technology" is a technology for analyzing intent and content from text data.
[0761] A "task management system" is a system that manages generated task tickets and allows users to access and modify them.
[0762] "Work support" refers to an assistive action that the robot performs based on a task ticket.
[0763] To implement this invention, the following components are required: smart glasses worn by workers in the factory, a server that performs data analysis, and a robot that supports the work.
[0764] Hardware and Software Configuration
[0765] Smart Glasses
[0766] Smart glasses, such as Google Glass Enterprise Edition, are equipped with built-in cameras and microphones to collect real-time video and audio data from the worker's perspective.
[0767] server
[0768] The server is a high-performance data analysis server, and can use Amazon Web Services' EC2, etc. The server uses the following software:
[0769] Speech recognition engine: Uses Google Cloud Speech-to-Text to convert voice data into text data.
[0770] Image recognition engine: Extracts task information from video data using OpenCV and TensorFlow.
[0771] Natural language processing technology: BERT is used to analyze task information from text data.
[0772] Emotion recognition engine: Affectiva is used to analyze emotional information from audio and video data.
[0773] Task management system: Use JIRA or similar to manage generated task tickets.
[0774] robot
[0775] The Universal Robots UR series is suitable as a work support robot for use in factories.
[0776] Examples of data collection
[0777] The smart glasses, which serve as the device, capture video from the worker's point of view in real time and collect audio data with a built-in microphone. For example, they can record video of a worker's hands as they assemble parts, along with audio of a worker saying, "This part doesn't fit properly."
[0778] Specific examples of data transmission
[0779] The device sends the collected video and audio data to a server via a Wi-Fi connection, where the data is compressed, encrypted, and securely transferred.
[0780] Specific examples of data analysis
[0781] The server converts the received voice data into text data using Google Cloud Speech-to-Text, and analyzes the speech, "The left screw won't tighten properly." Next, the video data is analyzed using OpenCV and TensorFlow to extract the screw installation position and the condition of the parts. At the same time, the emotion recognition engine Affectiva detects stress from the worker's voice.
[0782] Example of task ticket generation
[0783] The server automatically generates a task ticket based on the analyzed task information and emotion information. For example, a task ticket such as "Screw tightening", "Duties: Operations team", "Deadline: Immediately", and "Emotion: Stress" may be generated.
[0784] Task ticket notification and management example
[0785] The generated task ticket is registered in the task management system JIRA and notified to the work support robot. Based on this information, the robot displays a step-by-step video guide to the worker and supports the procedure. If the worker is feeling stressed, the robot automatically provides support for the work.
[0786] Examples of prompt statements
[0787] Invention details: A system that analyzes data collected by smart glasses on a server and generates task tickets.
[0788] Usage scenario: A robot that supports work in a factory operates based on video, audio, and emotional data.
[0789] Input data example: Video, audio, and emotion data sent from smart glasses.
[0790] Example of output task ticket: "Screw tightening" "Worker: Operations team" "Deadline: Immediately" "Emotion: Stress"
[0791] Thus, the present invention provides a system that allows for real-time work support and task management while taking into account the emotional state of the worker.
[0792] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0793] Step 1:
[0794] The device (smart glasses) collects video and audio data.
[0795] Input: Video from the user's perspective and surrounding audio data
[0796] Output: Real-time collected video and audio data
[0797] How it works: The smart glasses' built-in camera captures video from the user's point of view, and the built-in microphone collects audio, recording, for example, a worker assembling parts and their conversations.
[0798] Step 2:
[0799] The device sends the data to the server
[0800] Input: Collected video and audio data
[0801] Output: Video and audio data transferred to the server
[0802] How it works: The smart glasses send the data they collect to a server via Wi-Fi. The data is compressed and encrypted for secure transmission. For example, a video of the worker's hands and a voice message saying, "This part isn't fitting properly" are sent.
[0803] Step 3:
[0804] The server converts the audio data into text data.
[0805] Input: Audio data sent from the device
[0806] Output: Character data
[0807] Specific operation: The server converts voice data into text data using Google Cloud Speech-to-Text. For example, a speech saying "I can't tighten the left screw properly" is converted into text data.
[0808] Step 4:
[0809] The server analyzes the video data and extracts task information.
[0810] Input: Video data sent from the device
[0811] Output: Task information
[0812] How it works: The server uses OpenCV and TensorFlow to analyze the video data and extract information relevant to the task, such as the position of screws or the condition of parts.
[0813] Step 5:
[0814] The server analyzes the emotional data
[0815] Input: Audio and video data
[0816] Output: Emotion data
[0817] How it works: The server uses Affectiva to analyze audio and video data and determine the user's emotions. For example, it can detect that a worker is feeling stressed.
[0818] Step 6:
[0819] The server analyzes the task information and generates a task ticket.
[0820] Input: Text data, task information extracted from video, emotion data
[0821] Output: Task ticket
[0822] Specific operation: The server analyzes the task information using BERT and generates a task ticket based on that information. For example, a task ticket with the following content might be created: "Screw tightening," "Duty: Operations team," "Deadline: Immediate," and "Emotion: Stress."
[0823] Step 7:
[0824] The server registers the task ticket in the task management system and notifies it.
[0825] Input: Generated task ticket
[0826] Output: Task tickets registered in the task management system
[0827] Specific operation: The task ticket generated by the server is registered in the task management system (e.g. JIRA) and this information is then notified to the robot.
[0828] Step 8:
[0829] Robots provide support and assistance for tasks
[0830] Input: Task ticket notified from the server
[0831] Output: Assistance performed
[0832] Specific behavior: The robot performs work support and assistance based on the contents of the task ticket, for example, displaying a step-by-step video guide when a worker is feeling stressed or bringing parts for the worker.
[0833] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0834] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0835] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0836] [Third embodiment]
[0837] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0838] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0839] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0840] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0841] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0842] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0843] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0844] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0845] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0846] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0847] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0848] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0849] To implement this invention, we will build a system using smart glasses and a server connected to them. This system collects video and audio data from the user's point of view, analyzes it on the server, generates task tickets, and registers them in a task management system.
[0850] Program processing overview
[0851] Data collection
[0852] The device uses the smart glasses' built-in camera to capture real-time video from the user's perspective, and a built-in microphone to collect the user's conversation and surrounding sounds, allowing for accurate recording of task information during meetings and conversations.
[0853] Example: The device simultaneously records whiteboard content and conversations during meetings. When the user is looking at a PC screen, the device also captures on-screen task information.
[0854] Data transmission
[0855] The devices transmit the collected video and audio data in real time to a server over an internet connection, where the data is compressed for efficient transmission and encrypted to ensure the information is secure.
[0856] Example: A device uses a Wi-Fi connection to send video and audio data to a server. Data can be sent even when you are away from home.
[0857] Data analysis
[0858] The server inputs the received video data into an image recognition engine to extract important elements including task information. At the same time, it inputs the audio data into a voice recognition engine and converts it into text data. Then, using natural language processing technology, it analyzes the task information from this text data.
[0859] Example: The server converts audio data such as "Write a report and submit it by next Monday" into text and extracts the task elements "Write a report," "Submit," and "Next Monday." It also analyzes a whiteboard note from video data that reads "Deadline: May 20th" and recognizes the deadline.
[0860] Generate a Task Ticket
[0861] The server automatically generates a task ticket based on the parsed task information, which includes information such as the task summary, details, deadline, and assignee, allowing users to manage tasks accurately and quickly without any manual intervention.
[0862] Example: A task ticket with the following name will be automatically generated: "Create a report", "Submission deadline: next Monday", "Responsible person: User".
[0863] Ticket notifications and management
[0864] The server registers the generated task ticket in the task management system and has a notification means for notifying the user. The user receives the notification, accesses the task management system, checks the generated task ticket, and can correct or enter additional information as necessary.
[0865] Example: The server notifies the user that a "New Task Ticket: Create Report" has been created, and the user opens a task management app to check the contents of the ticket. If necessary, they can enter notes or supplementary information and manage the progress of the task.
[0866] This system allows users to efficiently manage various tasks that arise during meetings, conversations, and work, significantly reducing the burden and errors of manual ticket creation.
[0867] The processing flow will be explained below.
[0868] Step 1:
[0869] The device uses a built-in camera to capture video from the user's point of view in real time, including the contents of the whiteboard, presentation slides, and PC screen during the meeting.
[0870] Step 2:
[0871] The device uses a built-in microphone to collect audio data in real time, recording the user's conversations and the sounds of the surrounding area.
[0872] Step 3:
[0873] The devices transmit the collected video and audio data to a server over the Internet, where the data is pre-compressed and encrypted to ensure efficient and secure transmission.
[0874] Step 4:
[0875] The server inputs the received voice data into a speech recognition engine and converts the voice into text data, making the conversation available in text format.
[0876] Step 5:
[0877] The server inputs the received video data into an image recognition engine and extracts task-related information from the video, for example, identifying text written on a whiteboard or task information displayed on a PC screen.
[0878] Step 6:
[0879] The server uses natural language processing technology to analyze the converted text data and identify task information, such as the task content, deadline, and person in charge, from the conversation and video data.
[0880] Step 7:
[0881] The server automatically generates a task ticket based on the analyzed task information. The generated task ticket includes information such as the task summary, details, deadline, and person in charge.
[0882] Step 8:
[0883] The server registers the generated task ticket in the task management system, which allows tasks to be centrally managed and users to check them later.
[0884] Step 9:
[0885] The server will notify the user of the newly created task ticket, via email, push notification, or other methods.
[0886] Step 10:
[0887] The user receives a notification and accesses the task management system to view the generated task ticket, enter additional notes and details as needed, and manage the task's progress.
[0888] Example 1
[0889] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0890] Conventional task management systems require task information to be entered manually, a time-consuming process that poses a high risk of input errors. Furthermore, there was no way to efficiently manage tasks that arise during meetings or conversations, which meant there was a risk of important information being overlooked or tasks being left out.
[0891] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0892] In this invention, the server includes means for capturing video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for compressing and encrypting the transmitted data, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, and means for notifying the user of the generated task tickets. This makes it possible to reliably collect and analyze task information in real time even during meetings and conversations, and automatically generate and manage task tickets.
[0893] An "integrated camera" is a camera that is built into a device and has the ability to capture video from the user's perspective in real time.
[0894] An "internal microphone" is a microphone that is built into a device and has the ability to collect audio data in real time.
[0895] A "server" is a computer system on a network that receives collected data and performs analysis and task ticket generation.
[0896] "Real-time" is a time processing concept that means data is generated, transmitted, and analyzed almost immediately.
[0897] "Speech recognition technology" is a technology that analyzes voice data and converts the content into text data.
[0898] "Character data" is text information obtained from voice data using voice recognition technology.
[0899] "Image recognition technology" is a technology that extracts specific information and patterns from video data.
[0900] "Natural language processing technology" is a general term for algorithms and technologies that analyze text data and understand, classify, and extract its content.
[0901] A "task ticket" is a manageable unit of a task that includes information such as task details, deadlines, and people in charge.
[0902] "Task Management System" means software or platform that registers and tracks generated task tickets and assists users in managing their tasks.
[0903] "Compression" is the process of reducing the size of data, a technique used to improve transmission efficiency.
[0904] "Encryption" is a technology that converts data to protect it and prevent unauthorized access or tampering during transmission.
[0905] A "push notification" is a notification message sent in real time from a server to a user's device.
[0906] To implement this invention, it is necessary to build a system using smart glasses and a server connected to them. The present invention collects video and audio data from the user's perspective, analyzes it on the server, generates task tickets, and registers them in a task management system.
[0907] First, the user puts on the smart glasses. The smart glasses are equipped with a built-in camera and microphone, which are used to collect video and audio data in real time. Specifically, the camera captures the video the user is looking at, and the microphone collects the conversations of the user and those around them. Typical smart glasses products can be used as the specific hardware for this purpose.
[0908] The device then transmits the collected video and audio data in real time to a server. During this transmission process, the data is compressed and encrypted. An internet connection is used as the communication medium. At this stage, data security and efficient transfer are crucial.
[0909] The server inputs the received video data into an image recognition engine and extracts key elements, including task information. Software used can include Google Cloud Vision API for image recognition, Google Cloud Speech-to-Text for voice recognition, and NLP technologies such as spaCy and Transformers for natural language processing. For example, if voice data such as "Write a report and submit it by next Monday" is sent to the server, it is converted into text using voice recognition technology, and task elements such as "Write a report," "Submit," and "Next Monday" are extracted. Furthermore, a whiteboard note from the video data stating "Deadline: May 20th" is analyzed, and the deadline is recognized.
[0910] Next, the server automatically generates a task ticket based on the analyzed task information. This task ticket includes the task summary, details, deadline, person in charge, etc. For example, specific information such as "Create a report," "Submission deadline: next Monday," and "Person in charge: User" is written in the task ticket.
[0911] The generated task ticket is registered in the task management system by the server. Commonly used platforms such as Trello and Jira can be used as the task management system. The server notifies the user of the generated task ticket. Possible notification methods include push notification and email.
[0912] Finally, the user accesses the task management system based on the received notification and checks the generated task ticket. The user can enter corrections or additional information into the task ticket as necessary to manage the progress of the task.
[0913] Specific examples
[0914] For example, if a user is told during a meeting to "write a report and submit it by next Monday," the user's video and audio will be collected through smart glasses. The device will then send the data to a server, which will then use speech recognition and natural language processing to extract task information. This will automatically generate a task ticket in the task management system with the following information: "Create a report," "Submission deadline: next Monday," and "Responsible party: User."
[0915] Prompt Sentence Examples
[0916] "Interpret the notes on the whiteboard during the meeting and extract important task information. For example, generate a task ticket with the content 'Submit report due May 20th'."
[0917] This system makes it possible to efficiently manage the various tasks that arise during meetings and conversations, significantly reducing the effort required to manually enter tasks and the risk of entering them incorrectly.
[0918] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0919] Step 1:
[0920] The device uses the camera built into the smart glasses to capture video from the user's point of view in real time, while simultaneously using the built-in microphone to collect the user's conversation and surrounding sounds. The input data is video data and audio data, and the output data is the collected raw data.
[0921] Specific behavior:
[0922] The device activates the smart glasses camera.
[0923] The device begins collecting audio data using the built-in microphone.
[0924] When a user is looking at notes on a whiteboard during a meeting, the terminal records the contents of the whiteboard as video.
[0925] When a user speaks, the microphone in the smart glasses picks up the sound.
[0926] Step 2:
[0927] The terminal transmits collected video and audio data to the server in real time. The data is compressed for efficiency and encrypted for security. The input data is the collected raw data. The output data is the compressed and encrypted data.
[0928] Specific behavior:
[0929] The device compresses the video and audio data.
[0930] The device encrypts the data.
[0931] The device sends the data to the server through an internet connection.
[0932] The server decompresses and decrypts the received data.
[0933] Step 3:
[0934] The server inputs the received video data into an image recognition engine to extract important elements including task information. At the same time, it inputs the audio data into a speech recognition engine and converts it into text data. The input data is raw data that has been decompressed and decrypted. The output data is the analyzed task information.
[0935] Specific behavior:
[0936] The server sends the video data to an image recognition engine, for example, using the Google Cloud Vision API.
[0937] The server sends the voice data to a speech recognition engine and converts it into text data. For example, Google Cloud Speech-to-Text is used.
[0938] The server sends the converted text data to a natural language processing engine, which analyzes and extracts task information. For example, we use spaCy.
[0939] The server extracts task details such as "write a report," "submit," and "next Monday" from the voice data.
[0940] The server analyzes the whiteboard note "Deadline: May 20th" from the video data and recognizes the task deadline.
[0941] Step 4:
[0942] The server automatically generates a task ticket based on the parsed task information. The task ticket contains information such as the task summary, details, deadline, and person in charge. The input data is the parsed task information. The output data is the generated task ticket.
[0943] Specific behavior:
[0944] The server automatically generates a task ticket based on the analyzed task information.
[0945] The server writes information such as "Create report," "Submission deadline: next Monday," and "Responsible person: User" into the task ticket.
[0946] Step 5:
[0947] The server registers the generated task ticket in the task management system and notifies the user. The input data is the generated task ticket. The output data is the notified task ticket information.
[0948] Specific behavior:
[0949] The server registers the task ticket in a task management system, such as Trello or Jira.
[0950] The server notifies the user of the new task ticket creation via push notification or email.
[0951] A user accesses the task management system to view the new task ticket and enter notes or additional information as needed.
[0952] (Application example 1)
[0953] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0954] In modern factory environments, many tasks are still performed manually, and efficiency and accuracy are essential. However, workers spend a lot of time manually creating and managing task tickets, resulting in reduced productivity. Furthermore, there is no system in place to accurately record the various tasks that occur during work and instantly notify robots. This disrupts the smooth progress of work within the factory, making it difficult to address any errors or delays that arise.
[0955] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0956] In this invention, the server includes means for capturing video from a user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, means for notifying the generated task tickets to users, means for having workers wear smart glasses in a factory environment to collect task details, send the collected data to the server for analysis, and initiate task notification to factory robots, and means for supporting workers to receive work instructions in real time and perform tasks efficiently, thereby enabling automation and improved accuracy of task management within the factory.
[0957] "Built-in camera" refers to a camera module built into an apparatus or device that captures video from the user's perspective in real time.
[0958] "Built-in microphone" refers to a microphone built into an apparatus or device for collecting audio data in real time.
[0959] A "server" refers to a computer system that processes data over a network and has storage and analysis capabilities.
[0960] "Speech recognition technology" refers to technology that analyzes voice data and converts that voice into text data.
[0961] "Image recognition technology" refers to technology for analyzing video data and extracting specific information from it.
[0962] "Natural language processing technology" refers to technology that analyzes text data and extracts and understands information based on human language.
[0963] A "task ticket" is an electronic card or format that compiles information about a specific task, including the task name, details, deadline, and assignee.
[0964] A "task management system" refers to software for registering and managing generated task tickets.
[0965] "Smart glasses" refers to a wearable eyeglass-type device that has augmented reality and information display functions.
[0966] "Factory robot" refers to a robot used to automate production tasks in a factory.
[0967] "Worker" refers to a person who actually performs manual work or supervises in a factory or production site.
[0968] To implement this invention, a system using a server linked to smart glasses is constructed as follows.
[0969] Data collection
[0970] Device:
[0971] The user wears the smart glasses, and while working, the built-in camera captures video from the user's point of view in real time. The built-in microphone collects audio data. Examples of hardware used to collect this data include the built-in camera from Logitech and the microphone from Blue Yeti.
[0972] Specific examples
[0973] Workers on the factory assembly line wear smart glasses and perform their tasks while receiving verbal instructions on part installation procedures and important points to note, which then record the work procedures as visual and audio data.
[0974] Data transmission
[0975] Device:
[0976] The collected video and audio data is sent to a server via the Internet, where it is compressed and encrypted to ensure efficient and secure transmission.
[0977] Data analysis
[0978] server:
[0979] The server inputs the received video data into an image recognition engine (e.g., OpenCV, TensorFlow) to extract key elements including task information. It also inputs the audio data into a speech recognition engine (e.g., Google Speech-to-Text API) to convert it into text data. It then uses natural language processing technology (e.g., NLTK, spaCy) to analyze the text data for specific task information.
[0980] Specific examples
[0981] The server converts the voice data, "Tighten this bolt according to step 7," into text and extracts the task elements "bolt tightening task" and "step 7." It also analyzes and recognizes signs from the video data that include the instruction "inspect after task completion."
[0982] Task ticket generation and notification
[0983] server:
[0984] A task ticket is automatically generated based on the analyzed task information. The task ticket contains information such as the task summary, details, deadline, and person in charge. The generated task ticket is registered in the task management system, and its contents are notified to the user.
[0985] Specific examples
[0986] Task tickets for "Bolt tightening," "Step 7," and "Inspect after completion" are automatically generated and displayed on the worker's smart glasses. The same information is also sent to the factory robot, which automatically supports the necessary tasks.
[0987] Real-time support for work instructions
[0988] Terminals and robots:
[0989] The smart glasses help workers receive real-time work instructions to carry out their work efficiently. They also display the next steps and important points to note, and send instructions to factory robots.
[0990] Prompt Sentence Examples
[0991] If you're having a discussion in a conference room:
[0992] I need to write and submit the report by next Monday.
[0993] For factory assembly lines:
[0994] "When you assemble this part to the next block, you need to tighten the nuts."
[0995] These processes significantly improve task management and work efficiency within the factory, reducing the burden and errors of manual ticket creation.
[0996] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0997] Program processing flow
[0998] Step 1: Data collection
[0999] Device:
[1000] Using the built-in camera, video data is captured in real time from the user's perspective.
[1001] The built-in microphone is used to collect the user's voice data in real time.
[1002] Input: Video data, audio data
[1003] Output: Collected video and audio files
[1004] Step 2: Send data
[1005] Device:
[1006] The collected video and audio data is sent to a server via the Internet.
[1007] Data is compressed and encrypted for secure and efficient transfer.
[1008] Input: Collected video and audio files
[1009] Output: Data sent to the server
[1010] Step 3: Data analysis (audio)
[1011] server:
[1012] The voice data is input into a voice recognition engine, and the voice is converted into text data.
[1013] Use voice recognition technology (e.g., Google Speech-to-Text API).
[1014] Input: Audio data
[1015] Output: Text data
[1016] Step 4: Data analysis (video)
[1017] server:
[1018] The video data is input into an image recognition engine, and task information is extracted from the video.
[1019] Uses image recognition technology (e.g., OpenCV, TensorFlow).
[1020] Input: Video data
[1021] Output: Extracted task information
[1022] Step 5: Parse task information
[1023] server:
[1024] Text data is input into natural language processing technology to analyze task information.
[1025] Use natural language processing technology (e.g., NLTK, spaCy).
[1026] Input: Text data
[1027] Output: Parsed task information
[1028] Step 6: Generate a task ticket
[1029] server:
[1030] Based on the parsed task information, a task ticket is automatically generated.
[1031] The task information includes a summary of the task, details, deadlines, and the person in charge.
[1032] Input: Parsed task information
[1033] Output: Generated task ticket
[1034] Step 7: Task ticket registration and notification
[1035] server:
[1036] The generated task ticket is registered in the task management system and notified to the user.
[1037] Input: Generated task ticket
[1038] Output: Notify user and register with task management system
[1039] Step 8: Real-time support for work instructions
[1040] Terminals and robots:
[1041] Workers receive work instructions in real time and are supported to carry out their work efficiently.
[1042] The smart glasses display the next steps and precautions to take, and send instructions to factory robots.
[1043] Input: Task ticket information
[1044] Output: Work instructions for smart glasses and factory robots
[1045] Specific operation explanation
[1046] Step 1: The device uses the smart glasses' built-in camera and microphone to collect real-time video and audio from the worker's perspective. The worker explains how to install the parts as they go along, collecting the necessary data.
[1047] Step 2: The collected video and audio data is sent to a server via the internet. The data is compressed and encrypted for efficient and secure transmission.
[1048] Step 3: The server uses a speech recognition engine to convert the voice data into text data. For example, the voice saying "Tighten this bolt according to step 7" is converted into text data.
[1049] Step 4: The server uses a video recognition engine to extract task information from the video data. For example, it can read "Step 7" and "Check after inspection" from signs and labels in the video.
[1050] Step 5: The server uses natural language processing technology to analyze specific task information from the converted text data. For example, it extracts task elements such as "bolt tightening," "step 7," and "inspection after completion of work."
[1051] Step 6: The server automatically generates a task ticket based on the parsed task information. The task ticket includes the task name, details, deadline, and person in charge.
[1052] Step 7: The generated task ticket is registered in the task management system and simultaneously notified to the worker's smart glasses and factory robots. The worker can then view the new task through the smart glasses.
[1053] Step 8: Using smart glasses, workers receive real-time instructions on the next steps and important points to note, allowing them to work efficiently. Factory robots also receive instructions and automatically assist with the designated tasks.
[1054] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1055] To implement this invention, we will build a system using smart glasses and a server equipped with an emotion engine linked to the glasses. This system collects video and audio data from the user's perspective, analyzes it on the server, generates task tickets, performs emotion recognition, registers them in a task management system, and notifies the user.
[1056] Program processing overview
[1057] Data collection
[1058] The device uses the smart glasses' built-in camera to capture real-time video from the user's point of view, and a built-in microphone to collect the user's conversation and surrounding sounds. The emotion engine uses this data to analyze the user's emotions.
[1059] Example: The device simultaneously records the content of the whiteboard and conversations during a meeting. When the user is looking at the PC screen, the device also captures the task information on the screen. Furthermore, the device analyzes the user's emotions from facial expressions and tone of voice while they are speaking.
[1060] Data transmission
[1061] The device transmits the collected video, audio, and emotion analysis data to a server via the internet. This transmission uses an internet connection. The data is compressed to improve transmission efficiency and encrypted to ensure the security of the information.
[1062] Example: A device uses a Wi-Fi connection to send video and audio data, as well as emotion analysis data, to a server. Data can be sent even when the user is away from home.
[1063] Data analysis
[1064] The server inputs the received voice data into a voice recognition engine, which converts the voice into text data. It then inputs the video data into an image recognition engine, which extracts task-related information from the video. It then uses natural language processing technology to analyze the task information from the text data. The emotion engine also analyzes the user's emotions from the voice and video data.
[1065] Example: The server converts voice data such as "Write a report and submit it by next Monday" into text and extracts the task elements "Write a report," "Submit," and "Next Monday." It also analyzes a whiteboard note from video data that reads "Deadline: May 20th" to recognize the deadline. At the same time, the emotion engine analyzes whether the user's voice sounds tense or whether they are smiling.
[1066] Generate a Task Ticket
[1067] The server automatically generates a task ticket based on the parsed task information, which includes the task summary, details, deadline, assignee, and user sentiment information, allowing for more context-sensitive responses.
[1068] Example: A task ticket containing "Write a report", "Submission deadline: next Monday", "Responsible party: User", and "User emotion: Tension" is automatically generated.
[1069] Ticket notifications and management
[1070] The server registers the generated task ticket in the task management system and has a notification means to notify the user. The user receives the notification, accesses the task management system to check the generated task ticket, and can correct or add additional information as necessary. It is also possible to adjust the priority and response method of the task based on the emotional information.
[1071] Example: The server notifies the user that a "New Task Ticket: Create Report" has been created, and the user opens a task management app to check the contents of the ticket. If necessary, the user can enter notes or supplementary information and manage the progress of the task. Taking into account the user's emotional information, if the user is feeling nervous, the system will take appropriate measures such as lowering the priority of the task.
[1072] This system allows users to efficiently manage various tasks that arise during meetings, conversations, and work, and enables detailed task responses that take into account the user's emotional state, significantly reducing the burden and errors of manual task entry and improving users' overall work efficiency and satisfaction.
[1073] The processing flow will be explained below.
[1074] Step 1:
[1075] The device uses a built-in camera to capture real-time video from the user's point of view, including the contents of the whiteboard, presentation slides, and PC screen during the meeting.
[1076] Step 2:
[1077] The device uses a built-in microphone to collect voice data in real time, recording the user's conversation and surrounding sounds, and an emotion engine analyzes the tone and tempo of the voice to recognize the user's emotions.
[1078] Step 3:
[1079] The device then sends the collected video, audio, and emotion analysis data to a server via the internet. The data is compressed to improve transmission efficiency and encrypted to ensure the security of the information.
[1080] Step 4:
[1081] The server inputs the received voice data into a speech recognition engine and converts the voice into text data, making the conversation available in text format.
[1082] Step 5:
[1083] The server inputs the received video data into an image recognition engine and extracts task-related information from the video, for example, identifying text written on a whiteboard or task information displayed on a PC screen.
[1084] Step 6:
[1085] The server uses natural language processing technology to analyze task information from the converted text data. Specifically, it identifies the task content, deadline, and person in charge from the conversation and video data. In addition, an emotion engine analyzes the user's emotions from the audio and video data.
[1086] Step 7:
[1087] The server automatically generates a task ticket based on the analyzed task information and the user's emotion information. The generated task ticket includes the task summary, details, deadline, assignee, and user's emotion information.
[1088] Step 8:
[1089] The server registers the generated task ticket in the task management system. At that time, it is also possible to set the priority of the task based on the user's emotional information.
[1090] Step 9:
[1091] The server will notify the user of the newly created task ticket, via email, push notification, or other methods.
[1092] Step 10:
[1093] Users receive a notification and access their task management system to view the generated task ticket, enter additional notes and details as needed, manage the task's progress, and even adjust the task's priority based on sentiment information.
[1094] Example 2
[1095] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1096] In today's business environment, where a variety of tasks are being carried out simultaneously, manually inputting and managing tasks by users is laborious and prone to errors. Furthermore, task management that takes into account the user's emotional state is difficult, resulting in reduced work efficiency and satisfaction. Current task management systems are unable to automatically generate tasks in real time or manage tasks that take into account emotional information, so improvements are needed to address this issue.
[1097] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for compressing and encrypting the collected video data and audio data and transmitting them to the server, means for converting audio data into text data using voice recognition technology, means for extracting task information from video data using image recognition technology, means for analyzing task information from text data using natural language processing technology, means for analyzing user emotions from the collected video data and audio data using emotion analysis technology, means for automatically generating task tickets based on task information, means for adding emotion information to the generated task tickets, means for registering the generated task tickets in a task management system, and means for notifying the user of the generated task tickets. This enables efficient and accurate task management that also takes the user's emotional state into consideration.
[1098] A "built-in camera" is a photographing device built into a terminal to capture images from the user's point of view in real time.
[1099] A "built-in microphone" is a voice collection device built into a terminal for collecting voice data in real time.
[1100] "Collected Video Data" means visual information captured in real time by the built-in camera.
[1101] "Collected Audio Data" means audio information collected in real time by the built-in microphone.
[1102] "Means of compressing, encrypting and transmitting" refers to the technical means of compressing collected data for efficient transmission, encrypting it to ensure the security of communications and transmitting it to a server via the Internet.
[1103] "Speech recognition technology" is a technology that analyzes collected voice data and converts it into text data.
[1104] "Character data" is voice information analyzed by voice recognition technology expressed in text format.
[1105] "Image recognition technology" is a technology that analyzes collected video data and extracts task information from the video.
[1106] "Natural language processing technology" is a technology that analyzes task information from text data and understands meaning and intent.
[1107] "Emotion analysis technology" is a technology that analyzes collected video and audio data to recognize the user's emotional state.
[1108] "Task information" refers to information related to tasks and instructions extracted from the analyzed data.
[1109] A "task ticket" is an electronic ticket that includes task summary, details, deadline, person in charge, and emotion information, which is automatically generated based on analyzed task information.
[1110] A "task management system" is software for registering and managing generated task tickets.
[1111] "Means for notifying" refers to a method or tool for notifying a user of a generated task ticket.
[1112] To implement this invention, it is necessary to build a system using smart glasses and a server equipped with an emotion engine linked to the glasses. This system collects video and audio data from the user's perspective, analyzes it on the server to generate task tickets, performs emotion recognition, registers them in a task management system, and notifies the user.
[1113] Specifically, the system uses the following hardware and software:
[1114] 1. Hardware:
[1115] Smart glasses: Equipped with a built-in camera and microphone.
[1116] Built-in camera: A device for capturing video from the user's point of view in real time.
[1117] Built-in microphone: A device for collecting audio data in real time.
[1118] 2. Software:
[1119] Emotion engine: Software for analyzing the user's emotional state from video and audio data.
[1120] Speech recognition technology: A technology that converts collected voice data into text data.
[1121] Image recognition technology: A technology that extracts task information from video data.
[1122] Natural language processing technology: A technology that analyzes task information from text data.
[1123] Task management system: Software for managing generated task tickets.
[1124] Internet connection: A network environment for sending collected data to a server.
[1125] Here is an example of how the system works:
[1126] Data collection:
[1127] The smart glasses, which function as a device, use a built-in camera to capture images from the user's perspective in real time. The built-in microphone also collects the user's conversation and surrounding sounds. The emotion engine uses this data to analyze the user's emotions. For example, the device simultaneously records the contents of the whiteboard and conversation during a meeting, and captures the task information on the PC screen when the user is looking at it. Furthermore, the device analyzes emotions from the user's facial expressions and tone of voice when they are speaking.
[1128] Data transmission:
[1129] The device sends the collected video data, audio data, and emotion analysis data to a server via the Internet. The data is compressed and encrypted the moment it is collected, improving transmission efficiency and ensuring the security of the information. For example, the device can use a Wi-Fi connection to send video, audio, and emotion analysis data to a server. Similar data transmission is possible even when you are away from home.
[1130] Data Analysis:
[1131] The server inputs the received voice data into a voice recognition engine, which converts the voice into text data. Video data is also input into an image recognition engine, which extracts information related to the task. Natural language processing technology is then used to analyze the task information from this text data. At the same time, an emotion engine analyzes the user's emotions from the voice and video data. For example, if voice data such as "Write a report and submit it by next Monday" is converted into text, the task elements "Write a report," "Submit," and "Next Monday" are extracted. A note written on a whiteboard from the video data saying "Deadline: May 20th" is also analyzed, and the deadline is recognized. The emotion engine analyzes whether the user's voice sounds tense or if they are smiling.
[1132] Task Ticket Generation:
[1133] The server automatically generates a task ticket based on the analyzed task information. The task ticket includes the task summary, details, deadline, assignee, and emotion information, allowing for more context-specific responses. For example, a task ticket containing "Report creation," "Submission deadline: next Monday," "Assignee: User," and "User emotion: Tension" is automatically generated.
[1134] Ticket notifications and management:
[1135] The server registers the generated task ticket in the task management system and notifies the user. The user receives the notification, accesses the task management system to check the generated task ticket, and can make corrections or enter additional information as necessary. It is also possible to adjust the priority of the task and how to respond based on emotional information. For example, the server notifies the user that a "New Task Ticket: Create Report" has been generated, and the user opens the task management app to check the ticket contents, enter notes or supplementary information as necessary, and manage the progress of the task. Taking into account the user's emotional information, the system can respond by lowering the priority of the task if the user is feeling nervous.
[1136] Example prompt sentence:
[1137] "Please analyze the data recorded from the whiteboard contents and conversations during meetings acquired by smart glasses, and extract task information. For example, please capture content that includes information such as 'Deadline: May 20th', and generate a task ticket based on the deadline."
[1138] "Convert task information from user conversation data into text data, and automatically generate task tickets that include emotional data based on a sentiment analysis engine."
[1139] This allows users to efficiently manage a variety of tasks that arise during meetings, conversations, and work, and enables detailed task responses that take into account the user's emotional state. This significantly reduces the burden and errors of manual task entry, improving users' overall work efficiency and satisfaction.
[1140] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1141] Step 1: Data collection
[1142] The device uses a built-in camera to capture video data from the user's perspective in real time. The device's built-in microphone is used to collect surrounding sounds and the user's conversation. The emotion engine receives the collected video and audio data as input and analyzes the user's emotions. For example, it collects text written on a whiteboard and the content of conversations during a meeting, and also obtains emotional information from the user's facial expressions and tone of voice. As a result, collected video data, audio data, and emotional information are generated.
[1143] Step 2: Send data
[1144] The device sends the compressed and encrypted collected data (video data, audio data, and emotional information) to a server via the Internet. In this process, the collected data is sent as input, and the compressed and encrypted data is transferred to the server as output. For example, the data can be sent using a Wi-Fi connection or a mobile network to reach the server safely and efficiently.
[1145] Step 3: Data analysis
[1146] The server inputs the received voice data into a voice recognition engine and converts the user's speech into text data. In parallel, it inputs the video data into an image recognition engine and extracts task-related information from the video. Then, using natural language processing technology, it analyzes the task information from this text data. The emotion engine reanalyzes the user's emotions from the voice and video data. For example, voice data such as "Write and submit a report by next Monday" is converted into text data, and the task elements "Write a report," "Submit," and "Next Monday" are extracted. At the same time, the whiteboard note "Deadline: May 20th" is analyzed and recognized from the video data, and tension is detected from the user's tone of voice. As a result, task information and emotion information are generated.
[1147] Step 4: Task Ticket Generation
[1148] The server automatically generates a task ticket based on the analyzed task information and emotion information. The task ticket includes a summary of the task, details, deadline, person in charge, emotion information, etc. For example, a task ticket containing information such as "Report creation," "Submission deadline: next Monday," "Person in charge: User," and "User emotion: Tension" is created. As a result, a task ticket is generated that reflects the user's emotional state.
[1149] Step 5: Ticket notification and management
[1150] The server registers the generated task ticket in the task management system and notifies the user. The user receives the notification and accesses the task management system to check the generated task ticket. They can modify the task ticket or enter additional information as necessary. It is also possible to adjust the priority and response method of the task based on emotional information. For example, the user is notified that a "New Task Ticket: Create Report" has been created, and the user opens the task management app to check the contents of the ticket and enter notes or supplementary information as necessary. In addition, if the user is feeling nervous, the task priority may be lowered. As a result, efficient task management that takes emotional information into consideration is achieved.
[1151] (Application example 2)
[1152] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1153] Modern factories are becoming increasingly automated to increase productivity, but there are still many situations in which human workers are required. Furthermore, a worker's emotional state and stress level can directly affect work efficiency and quality, but there is still a lack of systems that can detect this in real time and provide work support. Therefore, there is a need for work support and task management that takes into account the emotions of workers.
[1154] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1155] In this invention, the server includes means for capturing video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, means for notifying the user of the generated task tickets, means for analyzing the user's emotional data and adjusting task priorities and response methods based on the emotional data, and means for a robot that receives a task ticket to automatically provide work support or assistance. This enables work support and task management in real time while taking into account the emotional state of the worker.
[1156] A "built-in camera" is a device that captures video from the user's perspective in real time.
[1157] A "built-in microphone" is a device for collecting audio data in real time.
[1158] A "server" is a computer system that analyzes received data and manages various tasks.
[1159] A "task ticket" is a digital ticket that is automatically generated based on analyzed task information and contains a task summary, details, deadline, responsible person, etc.
[1160] "Emotional data" refers to data relating to the emotional state of a user analyzed from their voice and video.
[1161] A "robot" is a mechanical device used to provide support and assistance in work within a factory.
[1162] "Speech recognition technology" is a technology for converting voice data into text data.
[1163] "Image recognition technology" is a technology for extracting useful information from video data.
[1164] "Natural language processing technology" is a technology for analyzing intent and content from text data.
[1165] A "task management system" is a system that manages generated task tickets and allows users to access and modify them.
[1166] "Work support" refers to an assistive action that the robot performs based on a task ticket.
[1167] To implement this invention, the following components are required: smart glasses worn by workers in the factory, a server that performs data analysis, and a robot that supports the work.
[1168] Hardware and Software Configuration
[1169] Smart Glasses
[1170] Smart glasses, such as Google Glass Enterprise Edition, are equipped with built-in cameras and microphones to collect real-time video and audio data from the worker's perspective.
[1171] server
[1172] The server is a high-performance data analysis server, and can use Amazon Web Services' EC2, etc. The server uses the following software:
[1173] Speech recognition engine: Uses Google Cloud Speech-to-Text to convert voice data into text data.
[1174] Image recognition engine: Extracts task information from video data using OpenCV and TensorFlow.
[1175] Natural language processing technology: BERT is used to analyze task information from text data.
[1176] Emotion recognition engine: Affectiva is used to analyze emotional information from audio and video data.
[1177] Task management system: Use JIRA or similar to manage generated task tickets.
[1178] robot
[1179] The Universal Robots UR series is suitable as a work support robot for use in factories.
[1180] Examples of data collection
[1181] The smart glasses, which serve as the device, capture video from the worker's point of view in real time and collect audio data with a built-in microphone. For example, they can record video of a worker's hands as they assemble parts, along with audio of a worker saying, "This part doesn't fit properly."
[1182] Specific examples of data transmission
[1183] The device sends the collected video and audio data to a server via a Wi-Fi connection, where the data is compressed, encrypted, and securely transferred.
[1184] Specific examples of data analysis
[1185] The server converts the received voice data into text data using Google Cloud Speech-to-Text, and analyzes the speech, "The left screw won't tighten properly." Next, the video data is analyzed using OpenCV and TensorFlow to extract the screw installation position and the condition of the parts. At the same time, the emotion recognition engine Affectiva detects stress from the worker's voice.
[1186] Example of task ticket generation
[1187] The server automatically generates a task ticket based on the analyzed task information and emotion information. For example, a task ticket such as "Screw tightening", "Duties: Operations team", "Deadline: Immediately", and "Emotion: Stress" may be generated.
[1188] Task ticket notification and management example
[1189] The generated task ticket is registered in the task management system JIRA and notified to the work support robot. Based on this information, the robot displays a step-by-step video guide to the worker and supports the procedure. If the worker is feeling stressed, the robot automatically provides support for the work.
[1190] Examples of prompt statements
[1191] Invention details: A system that analyzes data collected by smart glasses on a server and generates task tickets.
[1192] Usage scenario: A robot that supports work in a factory operates based on video, audio, and emotional data.
[1193] Input data example: Video, audio, and emotion data sent from smart glasses.
[1194] Example of output task ticket: "Screw tightening" "Worker: Operations team" "Deadline: Immediately" "Emotion: Stress"
[1195] Thus, the present invention provides a system that allows for real-time work support and task management while taking into account the emotional state of the worker.
[1196] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1197] Step 1:
[1198] The device (smart glasses) collects video and audio data.
[1199] Input: Video from the user's perspective and surrounding audio data
[1200] Output: Real-time collected video and audio data
[1201] How it works: The smart glasses' built-in camera captures video from the user's point of view, and the built-in microphone collects audio, recording, for example, a worker assembling parts and their conversations.
[1202] Step 2:
[1203] The device sends the data to the server
[1204] Input: Collected video and audio data
[1205] Output: Video and audio data transferred to the server
[1206] How it works: The smart glasses send the data they collect to a server via Wi-Fi. The data is compressed and encrypted for secure transmission. For example, a video of the worker's hands and a voice message saying, "This part isn't fitting properly" are sent.
[1207] Step 3:
[1208] The server converts the audio data into text data.
[1209] Input: Audio data sent from the device
[1210] Output: Character data
[1211] Specific operation: The server converts voice data into text data using Google Cloud Speech-to-Text. For example, a speech saying "I can't tighten the left screw properly" is converted into text data.
[1212] Step 4:
[1213] The server analyzes the video data and extracts task information.
[1214] Input: Video data sent from the device
[1215] Output: Task information
[1216] How it works: The server uses OpenCV and TensorFlow to analyze the video data and extract information relevant to the task, such as the position of screws or the condition of parts.
[1217] Step 5:
[1218] The server analyzes the emotional data
[1219] Input: Audio and video data
[1220] Output: Emotion data
[1221] How it works: The server uses Affectiva to analyze audio and video data and determine the user's emotions. For example, it can detect that a worker is feeling stressed.
[1222] Step 6:
[1223] The server analyzes the task information and generates a task ticket.
[1224] Input: Text data, task information extracted from video, emotion data
[1225] Output: Task ticket
[1226] Specific operation: The server analyzes the task information using BERT and generates a task ticket based on that information. For example, a task ticket with the following content might be created: "Screw tightening," "Duty: Operations team," "Deadline: Immediate," and "Emotion: Stress."
[1227] Step 7:
[1228] The server registers the task ticket in the task management system and notifies it.
[1229] Input: Generated task ticket
[1230] Output: Task tickets registered in the task management system
[1231] Specific operation: The task ticket generated by the server is registered in the task management system (e.g. JIRA) and this information is then notified to the robot.
[1232] Step 8:
[1233] Robots provide support and assistance for tasks
[1234] Input: Task ticket notified from the server
[1235] Output: Assistance performed
[1236] Specific behavior: The robot performs work support and assistance based on the contents of the task ticket, for example, displaying a step-by-step video guide when a worker is feeling stressed or bringing parts for the worker.
[1237] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1238] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1239] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1240] [Fourth embodiment]
[1241] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1242] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1243] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1244] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1245] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1246] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1247] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1248] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1249] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1250] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1251] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1252] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1253] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1254] To implement this invention, we will build a system using smart glasses and a server connected to them. This system collects video and audio data from the user's point of view, analyzes it on the server, generates task tickets, and registers them in a task management system.
[1255] Program processing overview
[1256] Data collection
[1257] The device uses the smart glasses' built-in camera to capture real-time video from the user's perspective, and a built-in microphone to collect the user's conversation and surrounding sounds, allowing for accurate recording of task information during meetings and conversations.
[1258] Example: The device simultaneously records whiteboard content and conversations during meetings. When the user is looking at a PC screen, the device also captures on-screen task information.
[1259] Data transmission
[1260] The devices transmit the collected video and audio data in real time to a server over an internet connection, where the data is compressed for efficient transmission and encrypted to ensure the information is secure.
[1261] Example: A device uses a Wi-Fi connection to send video and audio data to a server. Data can be sent even when you are away from home.
[1262] Data analysis
[1263] The server inputs the received video data into an image recognition engine to extract important elements including task information. At the same time, it inputs the audio data into a voice recognition engine and converts it into text data. Then, using natural language processing technology, it analyzes the task information from this text data.
[1264] Example: The server converts audio data such as "Write a report and submit it by next Monday" into text and extracts the task elements "Write a report," "Submit," and "Next Monday." It also analyzes a whiteboard note from video data that reads "Deadline: May 20th" and recognizes the deadline.
[1265] Generate a Task Ticket
[1266] The server automatically generates a task ticket based on the parsed task information, which includes information such as the task summary, details, deadline, and assignee, allowing users to manage tasks accurately and quickly without any manual intervention.
[1267] Example: A task ticket with the following name will be automatically generated: "Create a report", "Submission deadline: next Monday", "Responsible person: User".
[1268] Ticket notifications and management
[1269] The server registers the generated task ticket in the task management system and has a notification means for notifying the user. The user receives the notification, accesses the task management system, checks the generated task ticket, and can correct or enter additional information as necessary.
[1270] Example: The server notifies the user that a "New Task Ticket: Create Report" has been created, and the user opens a task management app to check the contents of the ticket. If necessary, they can enter notes or supplementary information and manage the progress of the task.
[1271] This system allows users to efficiently manage various tasks that arise during meetings, conversations, and work, significantly reducing the burden and errors of manual ticket creation.
[1272] The processing flow will be explained below.
[1273] Step 1:
[1274] The device uses a built-in camera to capture video from the user's point of view in real time, including the contents of the whiteboard, presentation slides, and PC screen during the meeting.
[1275] Step 2:
[1276] The device uses a built-in microphone to collect audio data in real time, recording the user's conversations and the sounds of the surrounding area.
[1277] Step 3:
[1278] The devices transmit the collected video and audio data to a server over the Internet, where the data is pre-compressed and encrypted to ensure efficient and secure transmission.
[1279] Step 4:
[1280] The server inputs the received voice data into a speech recognition engine and converts the voice into text data, making the conversation available in text format.
[1281] Step 5:
[1282] The server inputs the received video data into an image recognition engine and extracts task-related information from the video, for example, identifying text written on a whiteboard or task information displayed on a PC screen.
[1283] Step 6:
[1284] The server uses natural language processing technology to analyze the converted text data and identify task information, such as the task content, deadline, and person in charge, from the conversation and video data.
[1285] Step 7:
[1286] The server automatically generates a task ticket based on the analyzed task information. The generated task ticket includes information such as the task summary, details, deadline, and person in charge.
[1287] Step 8:
[1288] The server registers the generated task ticket in the task management system, which allows tasks to be centrally managed and users to check them later.
[1289] Step 9:
[1290] The server will notify the user of the newly created task ticket, via email, push notification, or other methods.
[1291] Step 10:
[1292] The user receives a notification and accesses the task management system to view the generated task ticket, enter additional notes and details as needed, and manage the task's progress.
[1293] Example 1
[1294] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1295] Conventional task management systems require task information to be entered manually, a time-consuming process that poses a high risk of input errors. Furthermore, there was no way to efficiently manage tasks that arise during meetings or conversations, which meant there was a risk of important information being overlooked or tasks being left out.
[1296] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1297] In this invention, the server includes means for capturing video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for compressing and encrypting the transmitted data, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, and means for notifying the user of the generated task tickets. This makes it possible to reliably collect and analyze task information in real time even during meetings and conversations, and automatically generate and manage task tickets.
[1298] An "integrated camera" is a camera that is built into a device and has the ability to capture video from the user's perspective in real time.
[1299] An "internal microphone" is a microphone that is built into a device and has the ability to collect audio data in real time.
[1300] A "server" is a computer system on a network that receives collected data and performs analysis and task ticket generation.
[1301] "Real-time" is a time processing concept that means data is generated, transmitted, and analyzed almost immediately.
[1302] "Speech recognition technology" is a technology that analyzes voice data and converts the content into text data.
[1303] "Character data" is text information obtained from voice data using voice recognition technology.
[1304] "Image recognition technology" is a technology that extracts specific information and patterns from video data.
[1305] "Natural language processing technology" is a general term for algorithms and technologies that analyze text data and understand, classify, and extract its content.
[1306] A "task ticket" is a manageable unit of a task that includes information such as task details, deadlines, and people in charge.
[1307] "Task Management System" means software or platform that registers and tracks generated task tickets and assists users in managing their tasks.
[1308] "Compression" is the process of reducing the size of data, a technique used to improve transmission efficiency.
[1309] "Encryption" is a technology that converts data to protect it and prevent unauthorized access or tampering during transmission.
[1310] A "push notification" is a notification message sent in real time from a server to a user's device.
[1311] To implement this invention, it is necessary to build a system using smart glasses and a server connected to them. The present invention collects video and audio data from the user's perspective, analyzes it on the server, generates task tickets, and registers them in a task management system.
[1312] First, the user puts on the smart glasses. The smart glasses are equipped with a built-in camera and microphone, which are used to collect video and audio data in real time. Specifically, the camera captures the video the user is looking at, and the microphone collects the conversations of the user and those around them. Typical smart glasses products can be used as the specific hardware for this purpose.
[1313] The device then transmits the collected video and audio data in real time to a server. During this transmission process, the data is compressed and encrypted. An internet connection is used as the communication medium. At this stage, data security and efficient transfer are crucial.
[1314] The server inputs the received video data into an image recognition engine and extracts key elements, including task information. Software used can include Google Cloud Vision API for image recognition, Google Cloud Speech-to-Text for voice recognition, and NLP technologies such as spaCy and Transformers for natural language processing. For example, if voice data such as "Write a report and submit it by next Monday" is sent to the server, it is converted into text using voice recognition technology, and task elements such as "Write a report," "Submit," and "Next Monday" are extracted. Furthermore, a whiteboard note from the video data stating "Deadline: May 20th" is analyzed, and the deadline is recognized.
[1315] Next, the server automatically generates a task ticket based on the analyzed task information. This task ticket includes the task summary, details, deadline, person in charge, etc. For example, specific information such as "Create a report," "Submission deadline: next Monday," and "Person in charge: User" is written in the task ticket.
[1316] The generated task ticket is registered in the task management system by the server. Commonly used platforms such as Trello and Jira can be used as the task management system. The server notifies the user of the generated task ticket. Possible notification methods include push notification and email.
[1317] Finally, the user accesses the task management system based on the received notification and checks the generated task ticket. The user can enter corrections or additional information into the task ticket as necessary to manage the progress of the task.
[1318] Specific examples
[1319] For example, if a user is told during a meeting to "write a report and submit it by next Monday," the user's video and audio will be collected through smart glasses. The device will then send the data to a server, which will then use speech recognition and natural language processing to extract task information. This will automatically generate a task ticket in the task management system with the following information: "Create a report," "Submission deadline: next Monday," and "Responsible party: User."
[1320] Prompt Sentence Examples
[1321] "Interpret the notes on the whiteboard during the meeting and extract important task information. For example, generate a task ticket with the content 'Submit report due May 20th'."
[1322] This system makes it possible to efficiently manage the various tasks that arise during meetings and conversations, significantly reducing the effort required to manually enter tasks and the risk of entering them incorrectly.
[1323] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1324] Step 1:
[1325] The device uses the camera built into the smart glasses to capture video from the user's point of view in real time, while simultaneously using the built-in microphone to collect the user's conversation and surrounding sounds. The input data is video data and audio data, and the output data is the collected raw data.
[1326] Specific behavior:
[1327] The device activates the smart glasses camera.
[1328] The device begins collecting audio data using the built-in microphone.
[1329] When a user is looking at notes on a whiteboard during a meeting, the terminal records the contents of the whiteboard as video.
[1330] When a user speaks, the microphone in the smart glasses picks up the sound.
[1331] Step 2:
[1332] The terminal transmits collected video and audio data to the server in real time. The data is compressed for efficiency and encrypted for security. The input data is the collected raw data. The output data is the compressed and encrypted data.
[1333] Specific behavior:
[1334] The device compresses the video and audio data.
[1335] The device encrypts the data.
[1336] The device sends the data to the server through an internet connection.
[1337] The server decompresses and decrypts the received data.
[1338] Step 3:
[1339] The server inputs the received video data into an image recognition engine to extract important elements including task information. At the same time, it inputs the audio data into a speech recognition engine and converts it into text data. The input data is raw data that has been decompressed and decrypted. The output data is the analyzed task information.
[1340] Specific behavior:
[1341] The server sends the video data to an image recognition engine, for example, using the Google Cloud Vision API.
[1342] The server sends the voice data to a speech recognition engine and converts it into text data. For example, Google Cloud Speech-to-Text is used.
[1343] The server sends the converted text data to a natural language processing engine, which analyzes and extracts task information. For example, we use spaCy.
[1344] The server extracts task details such as "write a report," "submit," and "next Monday" from the voice data.
[1345] The server analyzes the whiteboard note "Deadline: May 20th" from the video data and recognizes the task deadline.
[1346] Step 4:
[1347] The server automatically generates a task ticket based on the parsed task information. The task ticket contains information such as the task summary, details, deadline, and person in charge. The input data is the parsed task information. The output data is the generated task ticket.
[1348] Specific behavior:
[1349] The server automatically generates a task ticket based on the analyzed task information.
[1350] The server writes information such as "Create report," "Submission deadline: next Monday," and "Responsible person: User" into the task ticket.
[1351] Step 5:
[1352] The server registers the generated task ticket in the task management system and notifies the user. The input data is the generated task ticket. The output data is the notified task ticket information.
[1353] Specific behavior:
[1354] The server registers the task ticket in a task management system, such as Trello or Jira.
[1355] The server notifies the user of the new task ticket creation via push notification or email.
[1356] A user accesses the task management system to view the new task ticket and enter notes or additional information as needed.
[1357] (Application example 1)
[1358] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1359] In modern factory environments, many tasks are still performed manually, and efficiency and accuracy are essential. However, workers spend a lot of time manually creating and managing task tickets, resulting in reduced productivity. Furthermore, there is no system in place to accurately record the various tasks that occur during work and instantly notify robots. This disrupts the smooth progress of work within the factory, making it difficult to address any errors or delays that arise.
[1360] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1361] In this invention, the server includes means for capturing video from a user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, means for notifying the generated task tickets to users, means for having workers wear smart glasses in a factory environment to collect task details, send the collected data to the server for analysis, and initiate task notification to factory robots, and means for supporting workers to receive work instructions in real time and perform tasks efficiently, thereby enabling automation and improved accuracy of task management within the factory.
[1362] "Built-in camera" refers to a camera module built into an apparatus or device that captures video from the user's perspective in real time.
[1363] "Built-in microphone" refers to a microphone built into an apparatus or device for collecting audio data in real time.
[1364] A "server" refers to a computer system that processes data over a network and has storage and analysis capabilities.
[1365] "Speech recognition technology" refers to technology that analyzes voice data and converts that voice into text data.
[1366] "Image recognition technology" refers to technology for analyzing video data and extracting specific information from it.
[1367] "Natural language processing technology" refers to technology that analyzes text data and extracts and understands information based on human language.
[1368] A "task ticket" is an electronic card or format that compiles information about a specific task, including the task name, details, deadline, and assignee.
[1369] A "task management system" refers to software for registering and managing generated task tickets.
[1370] "Smart glasses" refers to a wearable eyeglass-type device that has augmented reality and information display functions.
[1371] "Factory robot" refers to a robot used to automate production tasks in a factory.
[1372] "Worker" refers to a person who actually performs manual work or supervises in a factory or production site.
[1373] To implement this invention, a system using a server linked to smart glasses is constructed as follows.
[1374] Data collection
[1375] Device:
[1376] The user wears the smart glasses, and while working, the built-in camera captures video from the user's point of view in real time. The built-in microphone collects audio data. Examples of hardware used to collect this data include the built-in camera from Logitech and the microphone from Blue Yeti.
[1377] Specific examples
[1378] Workers on the factory assembly line wear smart glasses and perform their tasks while receiving verbal instructions on part installation procedures and important points to note, which then record the work procedures as visual and audio data.
[1379] Data transmission
[1380] Device:
[1381] The collected video and audio data is sent to a server via the Internet, where it is compressed and encrypted to ensure efficient and secure transmission.
[1382] Data analysis
[1383] server:
[1384] The server inputs the received video data into an image recognition engine (e.g., OpenCV, TensorFlow) to extract key elements including task information. It also inputs the audio data into a speech recognition engine (e.g., Google Speech-to-Text API) to convert it into text data. It then uses natural language processing technology (e.g., NLTK, spaCy) to analyze the text data for specific task information.
[1385] Specific examples
[1386] The server converts the voice data, "Tighten this bolt according to step 7," into text and extracts the task elements "bolt tightening task" and "step 7." It also analyzes and recognizes signs from the video data that include the instruction "inspect after task completion."
[1387] Task ticket generation and notification
[1388] server:
[1389] A task ticket is automatically generated based on the analyzed task information. The task ticket contains information such as the task summary, details, deadline, and person in charge. The generated task ticket is registered in the task management system, and its contents are notified to the user.
[1390] Specific examples
[1391] Task tickets for "Bolt tightening," "Step 7," and "Inspect after completion" are automatically generated and displayed on the worker's smart glasses. The same information is also sent to the factory robot, which automatically supports the necessary tasks.
[1392] Real-time support for work instructions
[1393] Terminals and robots:
[1394] The smart glasses help workers receive real-time work instructions to carry out their work efficiently. They also display the next steps and important points to note, and send instructions to factory robots.
[1395] Prompt Sentence Examples
[1396] If you're having a discussion in a conference room:
[1397] I need to write and submit the report by next Monday.
[1398] For factory assembly lines:
[1399] "When you assemble this part to the next block, you need to tighten the nuts."
[1400] These processes significantly improve task management and work efficiency within the factory, reducing the burden and errors of manual ticket creation.
[1401] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1402] Program processing flow
[1403] Step 1: Data collection
[1404] Device:
[1405] Using the built-in camera, video data is captured in real time from the user's perspective.
[1406] The built-in microphone is used to collect the user's voice data in real time.
[1407] Input: Video data, audio data
[1408] Output: Collected video and audio files
[1409] Step 2: Send data
[1410] Device:
[1411] The collected video and audio data is sent to a server via the Internet.
[1412] Data is compressed and encrypted for secure and efficient transfer.
[1413] Input: Collected video and audio files
[1414] Output: Data sent to the server
[1415] Step 3: Data analysis (audio)
[1416] server:
[1417] The voice data is input into a voice recognition engine, and the voice is converted into text data.
[1418] Use voice recognition technology (e.g., Google Speech-to-Text API).
[1419] Input: Audio data
[1420] Output: Text data
[1421] Step 4: Data analysis (video)
[1422] server:
[1423] The video data is input into an image recognition engine, and task information is extracted from the video.
[1424] Uses image recognition technology (e.g., OpenCV, TensorFlow).
[1425] Input: Video data
[1426] Output: Extracted task information
[1427] Step 5: Parse task information
[1428] server:
[1429] Text data is input into natural language processing technology to analyze task information.
[1430] Use natural language processing technology (e.g., NLTK, spaCy).
[1431] Input: Text data
[1432] Output: Parsed task information
[1433] Step 6: Generate a task ticket
[1434] server:
[1435] Based on the parsed task information, a task ticket is automatically generated.
[1436] The task information includes a summary of the task, details, deadlines, and the person in charge.
[1437] Input: Parsed task information
[1438] Output: Generated task ticket
[1439] Step 7: Task ticket registration and notification
[1440] server:
[1441] The generated task ticket is registered in the task management system and notified to the user.
[1442] Input: Generated task ticket
[1443] Output: Notify user and register with task management system
[1444] Step 8: Real-time support for work instructions
[1445] Terminals and robots:
[1446] Workers receive work instructions in real time and are supported to carry out their work efficiently.
[1447] The smart glasses display the next steps and precautions to take, and send instructions to factory robots.
[1448] Input: Task ticket information
[1449] Output: Work instructions for smart glasses and factory robots
[1450] Specific operation explanation
[1451] Step 1: The device uses the smart glasses' built-in camera and microphone to collect real-time video and audio from the worker's perspective. The worker explains how to install the parts as they go along, collecting the necessary data.
[1452] Step 2: The collected video and audio data is sent to a server via the internet. The data is compressed and encrypted for efficient and secure transmission.
[1453] Step 3: The server uses a speech recognition engine to convert the voice data into text data. For example, the voice saying "Tighten this bolt according to step 7" is converted into text data.
[1454] Step 4: The server uses a video recognition engine to extract task information from the video data. For example, it can read "Step 7" and "Check after inspection" from signs and labels in the video.
[1455] Step 5: The server uses natural language processing technology to analyze specific task information from the converted text data. For example, it extracts task elements such as "bolt tightening," "step 7," and "inspection after completion of work."
[1456] Step 6: The server automatically generates a task ticket based on the parsed task information. The task ticket includes the task name, details, deadline, and person in charge.
[1457] Step 7: The generated task ticket is registered in the task management system and simultaneously notified to the worker's smart glasses and factory robots. The worker can then view the new task through the smart glasses.
[1458] Step 8: Using smart glasses, workers receive real-time instructions on the next steps and important points to note, allowing them to work efficiently. Factory robots also receive instructions and automatically assist with the designated tasks.
[1459] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1460] To implement this invention, we will build a system using smart glasses and a server equipped with an emotion engine linked to the glasses. This system collects video and audio data from the user's perspective, analyzes it on the server, generates task tickets, performs emotion recognition, registers them in a task management system, and notifies the user.
[1461] Program processing overview
[1462] Data collection
[1463] The device uses the smart glasses' built-in camera to capture real-time video from the user's point of view, and a built-in microphone to collect the user's conversation and surrounding sounds. The emotion engine uses this data to analyze the user's emotions.
[1464] Example: The device simultaneously records the content of the whiteboard and conversations during a meeting. When the user is looking at the PC screen, the device also captures the task information on the screen. Furthermore, the device analyzes the user's emotions from facial expressions and tone of voice while they are speaking.
[1465] Data transmission
[1466] The device transmits the collected video, audio, and emotion analysis data to a server via the internet. This transmission uses an internet connection. The data is compressed to improve transmission efficiency and encrypted to ensure the security of the information.
[1467] Example: A device uses a Wi-Fi connection to send video and audio data, as well as emotion analysis data, to a server. Data can be sent even when the user is away from home.
[1468] Data analysis
[1469] The server inputs the received voice data into a voice recognition engine, which converts the voice into text data. It then inputs the video data into an image recognition engine, which extracts task-related information from the video. It then uses natural language processing technology to analyze the task information from the text data. The emotion engine also analyzes the user's emotions from the voice and video data.
[1470] Example: The server converts voice data such as "Write a report and submit it by next Monday" into text and extracts the task elements "Write a report," "Submit," and "Next Monday." It also analyzes a whiteboard note from video data that reads "Deadline: May 20th" to recognize the deadline. At the same time, the emotion engine analyzes whether the user's voice sounds tense or whether they are smiling.
[1471] Generate a Task Ticket
[1472] The server automatically generates a task ticket based on the parsed task information, which includes the task summary, details, deadline, assignee, and user sentiment information, allowing for more context-sensitive responses.
[1473] Example: A task ticket containing "Write a report", "Submission deadline: next Monday", "Responsible party: User", and "User emotion: Tension" is automatically generated.
[1474] Ticket notifications and management
[1475] The server registers the generated task ticket in the task management system and has a notification means to notify the user. The user receives the notification, accesses the task management system to check the generated task ticket, and can correct or add additional information as necessary. It is also possible to adjust the priority and response method of the task based on the emotional information.
[1476] Example: The server notifies the user that a "New Task Ticket: Create Report" has been created, and the user opens a task management app to check the contents of the ticket. If necessary, the user can enter notes or supplementary information and manage the progress of the task. Taking into account the user's emotional information, if the user is feeling nervous, the system will take appropriate measures such as lowering the priority of the task.
[1477] This system allows users to efficiently manage various tasks that arise during meetings, conversations, and work, and enables detailed task responses that take into account the user's emotional state, significantly reducing the burden and errors of manual task entry and improving users' overall work efficiency and satisfaction.
[1478] The processing flow will be explained below.
[1479] Step 1:
[1480] The device uses a built-in camera to capture real-time video from the user's point of view, including the contents of the whiteboard, presentation slides, and PC screen during the meeting.
[1481] Step 2:
[1482] The device uses a built-in microphone to collect voice data in real time, recording the user's conversation and surrounding sounds, and an emotion engine analyzes the tone and tempo of the voice to recognize the user's emotions.
[1483] Step 3:
[1484] The device then sends the collected video, audio, and emotion analysis data to a server via the internet. The data is compressed to improve transmission efficiency and encrypted to ensure the security of the information.
[1485] Step 4:
[1486] The server inputs the received voice data into a speech recognition engine and converts the voice into text data, making the conversation available in text format.
[1487] Step 5:
[1488] The server inputs the received video data into an image recognition engine and extracts task-related information from the video, for example, identifying text written on a whiteboard or task information displayed on a PC screen.
[1489] Step 6:
[1490] The server uses natural language processing technology to analyze task information from the converted text data. Specifically, it identifies the task content, deadline, and person in charge from the conversation and video data. In addition, an emotion engine analyzes the user's emotions from the audio and video data.
[1491] Step 7:
[1492] The server automatically generates a task ticket based on the analyzed task information and the user's emotion information. The generated task ticket includes the task summary, details, deadline, assignee, and user's emotion information.
[1493] Step 8:
[1494] The server registers the generated task ticket in the task management system. At that time, it is also possible to set the priority of the task based on the user's emotional information.
[1495] Step 9:
[1496] The server will notify the user of the newly created task ticket, via email, push notification, or other methods.
[1497] Step 10:
[1498] Users receive a notification and access their task management system to view the generated task ticket, enter additional notes and details as needed, manage the task's progress, and even adjust the task's priority based on sentiment information.
[1499] Example 2
[1500] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1501] In today's business environment, where a variety of tasks are being carried out simultaneously, manually inputting and managing tasks by users is laborious and prone to errors. Furthermore, task management that takes into account the user's emotional state is difficult, resulting in reduced work efficiency and satisfaction. Current task management systems are unable to automatically generate tasks in real time or manage tasks that take into account emotional information, so improvements are needed to address this issue.
[1502] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for compressing and encrypting the collected video data and audio data and transmitting them to the server, means for converting audio data into text data using voice recognition technology, means for extracting task information from video data using image recognition technology, means for analyzing task information from text data using natural language processing technology, means for analyzing user emotions from the collected video data and audio data using emotion analysis technology, means for automatically generating task tickets based on task information, means for adding emotion information to the generated task tickets, means for registering the generated task tickets in a task management system, and means for notifying the user of the generated task tickets. This enables efficient and accurate task management that also takes the user's emotional state into consideration.
[1503] A "built-in camera" is a photographing device built into a terminal to capture images from the user's point of view in real time.
[1504] A "built-in microphone" is a voice collection device built into a terminal for collecting voice data in real time.
[1505] "Collected Video Data" means visual information captured in real time by the built-in camera.
[1506] "Collected Audio Data" means audio information collected in real time by the built-in microphone.
[1507] "Means of compressing, encrypting and transmitting" refers to the technical means of compressing collected data for efficient transmission, encrypting it to ensure the security of communications and transmitting it to a server via the Internet.
[1508] "Speech recognition technology" is a technology that analyzes collected voice data and converts it into text data.
[1509] "Character data" is voice information analyzed by voice recognition technology expressed in text format.
[1510] "Image recognition technology" is a technology that analyzes collected video data and extracts task information from the video.
[1511] "Natural language processing technology" is a technology that analyzes task information from text data and understands meaning and intent.
[1512] "Emotion analysis technology" is a technology that analyzes collected video and audio data to recognize the user's emotional state.
[1513] "Task information" refers to information related to tasks and instructions extracted from the analyzed data.
[1514] A "task ticket" is an electronic ticket that includes task summary, details, deadline, person in charge, and emotion information, which is automatically generated based on analyzed task information.
[1515] A "task management system" is software for registering and managing generated task tickets.
[1516] "Means for notifying" refers to a method or tool for notifying a user of a generated task ticket.
[1517] To implement this invention, it is necessary to build a system using smart glasses and a server equipped with an emotion engine linked to the glasses. This system collects video and audio data from the user's perspective, analyzes it on the server to generate task tickets, performs emotion recognition, registers them in a task management system, and notifies the user.
[1518] Specifically, the system uses the following hardware and software:
[1519] 1. Hardware:
[1520] Smart glasses: Equipped with a built-in camera and microphone.
[1521] Built-in camera: A device for capturing video from the user's point of view in real time.
[1522] Built-in microphone: A device for collecting audio data in real time.
[1523] 2. Software:
[1524] Emotion engine: Software for analyzing the user's emotional state from video and audio data.
[1525] Speech recognition technology: A technology that converts collected voice data into text data.
[1526] Image recognition technology: A technology that extracts task information from video data.
[1527] Natural language processing technology: A technology that analyzes task information from text data.
[1528] Task management system: Software for managing generated task tickets.
[1529] Internet connection: A network environment for sending collected data to a server.
[1530] Here is an example of how the system works:
[1531] Data collection:
[1532] The smart glasses, which function as a device, use a built-in camera to capture images from the user's perspective in real time. The built-in microphone also collects the user's conversation and surrounding sounds. The emotion engine uses this data to analyze the user's emotions. For example, the device simultaneously records the contents of the whiteboard and conversation during a meeting, and captures the task information on the PC screen when the user is looking at it. Furthermore, the device analyzes emotions from the user's facial expressions and tone of voice when they are speaking.
[1533] Data transmission:
[1534] The device sends the collected video data, audio data, and emotion analysis data to a server via the Internet. The data is compressed and encrypted the moment it is collected, improving transmission efficiency and ensuring the security of the information. For example, the device can use a Wi-Fi connection to send video, audio, and emotion analysis data to a server. Similar data transmission is possible even when you are away from home.
[1535] Data Analysis:
[1536] The server inputs the received voice data into a voice recognition engine, which converts the voice into text data. Video data is also input into an image recognition engine, which extracts information related to the task. Natural language processing technology is then used to analyze the task information from this text data. At the same time, an emotion engine analyzes the user's emotions from the voice and video data. For example, if voice data such as "Write a report and submit it by next Monday" is converted into text, the task elements "Write a report," "Submit," and "Next Monday" are extracted. A note written on a whiteboard from the video data saying "Deadline: May 20th" is also analyzed, and the deadline is recognized. The emotion engine analyzes whether the user's voice sounds tense or if they are smiling.
[1537] Task Ticket Generation:
[1538] The server automatically generates a task ticket based on the analyzed task information. The task ticket includes the task summary, details, deadline, assignee, and emotion information, allowing for more context-specific responses. For example, a task ticket containing "Report creation," "Submission deadline: next Monday," "Assignee: User," and "User emotion: Tension" is automatically generated.
[1539] Ticket notifications and management:
[1540] The server registers the generated task ticket in the task management system and notifies the user. The user receives the notification, accesses the task management system to check the generated task ticket, and can make corrections or enter additional information as necessary. It is also possible to adjust the priority of the task and how to respond based on emotional information. For example, the server notifies the user that a "New Task Ticket: Create Report" has been generated, and the user opens the task management app to check the ticket contents, enter notes or supplementary information as necessary, and manage the progress of the task. Taking into account the user's emotional information, the system can respond by lowering the priority of the task if the user is feeling nervous.
[1541] Example prompt sentence:
[1542] "Please analyze the data recorded from the whiteboard contents and conversations during meetings acquired by smart glasses, and extract task information. For example, please capture content that includes information such as 'Deadline: May 20th', and generate a task ticket based on the deadline."
[1543] "Convert task information from user conversation data into text data, and automatically generate task tickets that include emotional data based on a sentiment analysis engine."
[1544] This allows users to efficiently manage a variety of tasks that arise during meetings, conversations, and work, and enables detailed task responses that take into account the user's emotional state. This significantly reduces the burden and errors of manual task entry, improving users' overall work efficiency and satisfaction.
[1545] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1546] Step 1: Data collection
[1547] The device uses a built-in camera to capture video data from the user's perspective in real time. The device's built-in microphone is used to collect surrounding sounds and the user's conversation. The emotion engine receives the collected video and audio data as input and analyzes the user's emotions. For example, it collects text written on a whiteboard and the content of conversations during a meeting, and also obtains emotional information from the user's facial expressions and tone of voice. As a result, collected video data, audio data, and emotional information are generated.
[1548] Step 2: Send data
[1549] The device sends the compressed and encrypted collected data (video data, audio data, and emotional information) to a server via the Internet. In this process, the collected data is sent as input, and the compressed and encrypted data is transferred to the server as output. For example, the data can be sent using a Wi-Fi connection or a mobile network to reach the server safely and efficiently.
[1550] Step 3: Data analysis
[1551] The server inputs the received voice data into a voice recognition engine and converts the user's speech into text data. In parallel, it inputs the video data into an image recognition engine and extracts task-related information from the video. Then, using natural language processing technology, it analyzes the task information from this text data. The emotion engine reanalyzes the user's emotions from the voice and video data. For example, voice data such as "Write and submit a report by next Monday" is converted into text data, and the task elements "Write a report," "Submit," and "Next Monday" are extracted. At the same time, the whiteboard note "Deadline: May 20th" is analyzed and recognized from the video data, and tension is detected from the user's tone of voice. As a result, task information and emotion information are generated.
[1552] Step 4: Task Ticket Generation
[1553] The server automatically generates a task ticket based on the analyzed task information and emotion information. The task ticket includes a summary of the task, details, deadline, person in charge, emotion information, etc. For example, a task ticket containing information such as "Report creation," "Submission deadline: next Monday," "Person in charge: User," and "User emotion: Tension" is created. As a result, a task ticket is generated that reflects the user's emotional state.
[1554] Step 5: Ticket notification and management
[1555] The server registers the generated task ticket in the task management system and notifies the user. The user receives the notification and accesses the task management system to check the generated task ticket. They can modify the task ticket or enter additional information as necessary. It is also possible to adjust the priority and response method of the task based on emotional information. For example, the user is notified that a "New Task Ticket: Create Report" has been created, and the user opens the task management app to check the contents of the ticket and enter notes or supplementary information as necessary. In addition, if the user is feeling nervous, the task priority may be lowered. As a result, efficient task management that takes emotional information into consideration is achieved.
[1556] (Application example 2)
[1557] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1558] Modern factories are becoming increasingly automated to increase productivity, but there are still many situations in which human workers are required. Furthermore, a worker's emotional state and stress level can directly affect work efficiency and quality, but there is still a lack of systems that can detect this in real time and provide work support. Therefore, there is a need for work support and task management that takes into account the emotions of workers.
[1559] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1560] In this invention, the server includes means for capturing video from the user's viewpoint in real time using a built-in camera, means for collecting audio data in real time using a built-in microphone, means for transmitting the collected video and audio data to the server, means for converting the audio data into text data using voice recognition technology, means for extracting task information from the video data using image recognition technology, means for analyzing the task information from the text data using natural language processing technology, means for automatically generating task tickets based on the task information, means for registering the generated task tickets in a task management system, means for notifying the user of the generated task tickets, means for analyzing the user's emotional data and adjusting task priorities and response methods based on the emotional data, and means for a robot that receives a task ticket to automatically provide work support or assistance. This enables work support and task management in real time while taking into account the emotional state of the worker.
[1561] A "built-in camera" is a device that captures video from the user's perspective in real time.
[1562] A "built-in microphone" is a device for collecting audio data in real time.
[1563] A "server" is a computer system that analyzes received data and manages various tasks.
[1564] A "task ticket" is a digital ticket that is automatically generated based on analyzed task information and contains a task summary, details, deadline, responsible person, etc.
[1565] "Emotional data" refers to data relating to the emotional state of a user analyzed from their voice and video.
[1566] A "robot" is a mechanical device used to provide support and assistance in work within a factory.
[1567] "Speech recognition technology" is a technology for converting voice data into text data.
[1568] "Image recognition technology" is a technology for extracting useful information from video data.
[1569] "Natural language processing technology" is a technology for analyzing intent and content from text data.
[1570] A "task management system" is a system that manages generated task tickets and allows users to access and modify them.
[1571] "Work support" refers to an assistive action that the robot performs based on a task ticket.
[1572] To implement this invention, the following components are required: smart glasses worn by workers in the factory, a server that performs data analysis, and a robot that supports the work.
[1573] Hardware and Software Configuration
[1574] Smart Glasses
[1575] Smart glasses, such as Google Glass Enterprise Edition, are equipped with built-in cameras and microphones to collect real-time video and audio data from the worker's perspective.
[1576] server
[1577] The server is a high-performance data analysis server, and can use Amazon Web Services' EC2, etc. The server uses the following software:
[1578] Speech recognition engine: Uses Google Cloud Speech-to-Text to convert voice data into text data.
[1579] Image recognition engine: Extracts task information from video data using OpenCV and TensorFlow.
[1580] Natural language processing technology: BERT is used to analyze task information from text data.
[1581] Emotion recognition engine: Affectiva is used to analyze emotional information from audio and video data.
[1582] Task management system: Use JIRA or similar to manage generated task tickets.
[1583] robot
[1584] The Universal Robots UR series is suitable as a work support robot for use in factories.
[1585] Examples of data collection
[1586] The smart glasses, which serve as the device, capture video from the worker's point of view in real time and collect audio data with a built-in microphone. For example, they can record video of a worker's hands as they assemble parts, along with audio of a worker saying, "This part doesn't fit properly."
[1587] Specific examples of data transmission
[1588] The device sends the collected video and audio data to a server via a Wi-Fi connection, where the data is compressed, encrypted, and securely transferred.
[1589] Specific examples of data analysis
[1590] The server converts the received voice data into text data using Google Cloud Speech-to-Text, and analyzes the speech, "The left screw won't tighten properly." Next, the video data is analyzed using OpenCV and TensorFlow to extract the screw installation position and the condition of the parts. At the same time, the emotion recognition engine Affectiva detects stress from the worker's voice.
[1591] Example of task ticket generation
[1592] The server automatically generates a task ticket based on the analyzed task information and emotion information. For example, a task ticket such as "Screw tightening", "Duties: Operations team", "Deadline: Immediately", and "Emotion: Stress" may be generated.
[1593] Task ticket notification and management example
[1594] The generated task ticket is registered in the task management system JIRA and notified to the work support robot. Based on this information, the robot displays a step-by-step video guide to the worker and supports the procedure. If the worker is feeling stressed, the robot automatically provides support for the work.
[1595] Examples of prompt statements
[1596] Invention details: A system that analyzes data collected by smart glasses on a server and generates task tickets.
[1597] Usage scenario: A robot that supports work in a factory operates based on video, audio, and emotional data.
[1598] Input data example: Video, audio, and emotion data sent from smart glasses.
[1599] Example of output task ticket: "Screw tightening" "Worker: Operations team" "Deadline: Immediately" "Emotion: Stress"
[1600] Thus, the present invention provides a system that allows for real-time work support and task management while taking into account the emotional state of the worker.
[1601] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1602] Step 1:
[1603] The device (smart glasses) collects video and audio data.
[1604] Input: Video from the user's perspective and surrounding audio data
[1605] Output: Real-time collected video and audio data
[1606] How it works: The smart glasses' built-in camera captures video from the user's point of view, and the built-in microphone collects audio, recording, for example, a worker assembling parts and their conversations.
[1607] Step 2:
[1608] The device sends the data to the server
[1609] Input: Collected video and audio data
[1610] Output: Video and audio data transferred to the server
[1611] How it works: The smart glasses send the data they collect to a server via Wi-Fi. The data is compressed and encrypted for secure transmission. For example, a video of the worker's hands and a voice message saying, "This part isn't fitting properly" are sent.
[1612] Step 3:
[1613] The server converts the audio data into text data.
[1614] Input: Audio data sent from the device
[1615] Output: Character data
[1616] Specific operation: The server converts voice data into text data using Google Cloud Speech-to-Text. For example, a speech saying "I can't tighten the left screw properly" is converted into text data.
[1617] Step 4:
[1618] The server analyzes the video data and extracts task information.
[1619] Input: Video data sent from the device
[1620] Output: Task information
[1621] How it works: The server uses OpenCV and TensorFlow to analyze the video data and extract information relevant to the task, such as the position of screws or the condition of parts.
[1622] Step 5:
[1623] The server analyzes the emotional data
[1624] Input: Audio and video data
[1625] Output: Emotion data
[1626] How it works: The server uses Affectiva to analyze audio and video data and determine the user's emotions. For example, it can detect that a worker is feeling stressed.
[1627] Step 6:
[1628] The server analyzes the task information and generates a task ticket.
[1629] Input: Text data, task information extracted from video, emotion data
[1630] Output: Task ticket
[1631] Specific operation: The server analyzes the task information using BERT and generates a task ticket based on that information. For example, a task ticket with the following content might be created: "Screw tightening," "Duty: Operations team," "Deadline: Immediate," and "Emotion: Stress."
[1632] Step 7:
[1633] The server registers the task ticket in the task management system and notifies it.
[1634] Input: Generated task ticket
[1635] Output: Task tickets registered in the task management system
[1636] Specific operation: The task ticket generated by the server is registered in the task management system (e.g. JIRA) and this information is then notified to the robot.
[1637] Step 8:
[1638] Robots provide support and assistance for tasks
[1639] Input: Task ticket notified from the server
[1640] Output: Assistance performed
[1641] Specific behavior: The robot performs work support and assistance based on the contents of the task ticket, for example, displaying a step-by-step video guide when a worker is feeling stressed or bringing parts for the worker.
[1642] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1643] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1644] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1645] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1646] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1647] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1648] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1649] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1650] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1651] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1652] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1653] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1654] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1655] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1656] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1657] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1658] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1659] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1660] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1661] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1662] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1663] The following is further disclosed regarding the above embodiment.
[1664] (Claim 1)
[1665] a means for capturing video from a user's point of view in real time using an integrated camera;
[1666] means for collecting audio data in real time using an internal microphone;
[1667] means for transmitting the collected video data and audio data to a server;
[1668] A means for converting voice data into text data using voice recognition technology;
[1669] A means for extracting task information from video data using image recognition technology;
[1670] A means for analyzing task information from character data using natural language processing technology;
[1671] A means for automatically generating task tickets based on task information;
[1672] A means for registering the generated task ticket in a task management system;
[1673] A way to notify users of generated task tickets
[1674] A system including:
[1675] (Claim 2)
[1676] 2. The system according to claim 1, further comprising means for transmitting the collected video data and audio data to a server via the Internet.
[1677] (Claim 3)
[1678] 10. The system according to claim 1, further comprising means for simultaneously analyzing task information using voice recognition technology and image recognition technology.
[1679] "Example 1"
[1680] (Claim 1)
[1681] a means for capturing video from a user's point of view in real time using an integrated camera;
[1682] means for collecting audio data in real time using an internal microphone;
[1683] means for transmitting the collected video data and audio data to a server;
[1684] means for compressing and encrypting the transmitted data;
[1685] A means for converting voice data into text data using voice recognition technology;
[1686] A means for extracting task information from video data using image recognition technology;
[1687] A means for analyzing task information from character data using natural language processing technology;
[1688] A means for automatically generating task tickets based on task information;
[1689] A means for registering the generated task ticket in a task management system;
[1690] A way to notify users of generated task tickets
[1691] A system including:
[1692] (Claim 2)
[1693] 2. The system according to claim 1, further comprising means for transmitting the collected video data and audio data to a server via the Internet.
[1694] (Claim 3)
[1695] 10. The system according to claim 1, further comprising means for simultaneously analyzing task information using voice recognition technology and image recognition technology.
[1696] "Application Example 1"
[1697] (Claim 1)
[1698] a means for capturing video from a user's point of view in real time using an integrated camera;
[1699] means for collecting audio data in real time using an internal microphone;
[1700] means for transmitting the collected video data and audio data to a server;
[1701] A means for converting voice data into text data using voice recognition technology;
[1702] A means for extracting task information from video data using image recognition technology;
[1703] A means for analyzing task information from character data using natural language processing technology;
[1704] A means for automatically generating task tickets based on task information;
[1705] A means for registering the generated task ticket in a task management system;
[1706] A way to notify users of generated task tickets
[1707] A system including:
[1708] In a factory environment, workers wear smart glasses to collect information about their work.
[1709] A means for transmitting the collected data to a server, analyzing it, and initiating task notification to the factory robot;
[1710] A means of supporting workers to carry out work efficiently while receiving work instructions in real time
[1711] A system including:
[1712] (Claim 2)
[1713] 2. The system according to claim 1, further comprising means for transmitting the collected video data and audio data to a server via the Internet.
[1714] (Claim 3)
[1715] 10. The system according to claim 1, further comprising means for simultaneously analyzing task information using voice recognition technology and image recognition technology.
[1716] "Example 2: Combining Emotion Engines"
[1717] (Claim 1)
[1718] a means for acquiring video from a user's point of view in real time using an integrated camera;
[1719] means for collecting audio data in real time using an internal microphone;
[1720] A means for compressing and encrypting the collected video and audio data and sending it to a server;
[1721] A means for converting voice data into text data using voice recognition technology;
[1722] A means for extracting task information from video data using image recognition technology;
[1723] A means for analyzing task information from character data using natural language processing technology;
[1724] A means for analyzing user emotions from the collected video data and audio data using emotion analysis technology;
[1725] A means for automatically generating task tickets based on task information;
[1726] A means for adding emotional information to the generated task ticket;
[1727] A means for registering the generated task ticket in a task management system;
[1728] A method for notifying users of generated task tickets
[1729] A system including:
[1730] (Claim 2)
[1731] 2. The system according to claim 1, further comprising means for transmitting the collected video data and audio data to a server via the Internet.
[1732] (Claim 3)
[1733] 2. The system according to claim 1, further comprising means for simultaneously analyzing task information and emotion information using voice recognition technology and image recognition technology.
[1734] "Application example 2 when combining emotion engines"
[1735] (Claim 1)
[1736] a means for capturing video from a user's point of view in real time using an integrated camera;
[1737] means for collecting audio data in real time using an internal microphone;
[1738] means for transmitting the collected video data and audio data to a server;
[1739] A means for converting voice data into text data using voice recognition technology;
[1740] A means for extracting task information from video data using image recognition technology;
[1741] A means for analyzing task information from character data using natural language processing technology;
[1742] A means for automatically generating task tickets based on task information;
[1743] A means for registering the generated task ticket in a task management system;
[1744] A means to notify users of the generated task ticket;
[1745] A method for analyzing user emotion data and adjusting task priorities and response methods based on those emotions;
[1746] A means for robots that receive task tickets to automatically provide work support and assistance
[1747] A system including:
[1748] (Claim 2)
[1749] 2. The system according to claim 1, further comprising means for transmitting the collected video data and audio data to a server via the Internet.
[1750] (Claim 3)
[1751] 10. The system according to claim 1, further comprising means for simultaneously analyzing task information using voice recognition technology and image recognition technology. [Explanation of symbols]
[1752] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for capturing video from a user's point of view in real time using an integrated camera; means for collecting audio data in real time using an internal microphone; means for transmitting the collected video data and audio data to a server; A means for converting voice data into text data using voice recognition technology; A means for extracting task information from video data using image recognition technology; A means for analyzing task information from character data using natural language processing technology; A means for automatically generating task tickets based on task information; A means for registering the generated task ticket in a task management system; A way to notify users of generated task tickets A system including:
2. 2. The system according to claim 1, further comprising means for transmitting the collected video data and audio data to a server via the Internet.
3. 2. The system according to claim 1, further comprising means for simultaneously analyzing task information using voice recognition technology and image recognition technology.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A