system
The system addresses the inefficiency of manual manual creation by using AI to generate clear, interactive manuals from video data, improving user comprehension and support through automated processes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Existing manual creation processes require significant manual effort and are difficult for beginners to understand, necessitating a more efficient and user-friendly solution.
A system comprising an analysis unit, generation unit, and support unit that uses generative AI to automatically create easy-to-understand manuals from video data, incorporating images and explanatory text, and provides a chatbot for question answering.
Automatically generates clear and efficient manuals from video content, reducing manual effort and enhancing user understanding through interactive elements and real-time support.
Smart Images

Figure 2026073293000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the prior art, there is a problem that a lot of man-hours are required for manual creation and it is difficult for beginners to understand.
[0005] The system according to the embodiment aims to automatically create an easy-to-understand manual based on video data.
Means for Solving the Problems
[0006] The system according to this embodiment comprises an analysis unit, a generation unit, a creation unit, and a support unit. The analysis unit analyzes a video. The generation unit generates images and explanatory text based on the content of the video analyzed by the analysis unit. The creation unit creates a manual based on the images and explanatory text generated by the generation unit. The support unit includes a chatbot that responds to questions about the manual created by the creation unit. [Effects of the Invention]
[0007] The system according to this embodiment can automatically create easy-to-understand manuals based on video data. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The manual creation system according to an embodiment of the present invention is a system that automatically creates a manual based on video footage. This manual creation system uses a generative AI to generate images and text, instantly creating a manual that is easy for anyone to understand. Furthermore, by providing a chatbot that answers questions about the manual, a more detailed understanding can be obtained. The manual creation system according to an embodiment of the present invention is a system that automatically creates a manual based on video footage. This system uses a generative AI to generate images and text, instantly creating a manual that is easy for anyone to understand. Furthermore, by providing a chatbot that answers questions about the manual, a more detailed understanding can be obtained. First, the user shoots a video. This video is input into the generative AI. The generative AI analyzes the content of the video and generates images and explanatory text for each step. For example, if a video of the machine operation procedure is input into the generative AI, the generative AI generates images and explanatory text for each operation step. Next, the manual is automatically created based on the images and explanatory text generated by the generative AI. The generated manual is structured to be easy for first-time users to understand. For example, the images and explanatory text for each step are arranged in order, so that the operation procedure can be understood at a glance. Furthermore, the generated manuals include a chatbot. When a user asks a question about the manual's content, the chatbot provides an answer. For example, if a user asks about a specific operating procedure, the chatbot will provide a detailed explanation of that procedure. This system significantly reduces the effort required to create manuals and simplifies version control. Because the generation AI automatically creates the manuals, the latest version is always provided. Also, because a chatbot is included, users can get answers to any questions they may have immediately. For example, in the manufacturing industry, when creating manuals for the operation of new machinery, it was previously necessary to create the manuals manually. However, by using this service, manuals can be automatically created simply by recording a video. This significantly improves work efficiency, and even first-time users can easily understand the operating procedures.This allows the manual creation system to efficiently provide easy-to-understand manuals by analyzing videos, generating images and explanatory text, creating manuals, and providing support through a chatbot.
[0029] The manual creation system according to this embodiment comprises an analysis unit, a generation unit, a creation unit, and a support unit. The analysis unit analyzes a video. The analysis unit analyzes the content of the video using, for example, a generation AI. The generation AI can analyze each frame of the video and detect specific objects or actions. For example, the analysis unit uses the generation AI to detect a specific operation procedure in the video and generates an image and explanatory text corresponding to that procedure. The generation unit generates images and explanatory text based on the content of the video analyzed by the analysis unit. The generation unit generates images and explanatory text for each step using, for example, a generation AI. The generation AI can generate images and explanatory text for each step based on the content of the video. For example, the generation unit uses the generation AI to generate images and explanatory text corresponding to a specific operation procedure in the video. The creation unit creates a manual based on the images and explanatory text generated by the generation unit. The creation unit creates a manual based on the generated images and explanatory text using, for example, a generation AI. The generation AI can automatically determine the structure and format of the manual based on the generated images and explanatory text. For example, the creation unit uses a generation AI to arrange images and explanatory text for each step in a sequential manner, creating a manual that allows users to understand the operating procedures at a glance. The support unit provides a chatbot that answers questions about the manual created by the creation unit. For example, the support unit uses AI to provide answers when a user asks a question about the contents of the manual. The AI can provide appropriate answers to the user's questions. For example, the support unit uses AI to provide detailed explanations in response to questions about specific operating procedures. Thus, the manual creation system according to this embodiment can efficiently provide easy-to-understand manuals by analyzing videos, generating images and explanatory text, creating manuals, and providing support through a chatbot.
[0030] The analysis unit analyzes the video. For example, the analysis unit uses a generative AI to analyze the content of the video. The generative AI can analyze each frame of the video and detect specific objects or actions. Specifically, the generative AI uses deep learning technology to analyze each frame in the video individually, and while considering the continuity between frames, it detects specific operation procedures and important actions. For example, it recognizes hand movements, tool usage, and changes on the screen in the video with high accuracy and extracts them as operation procedures. Furthermore, the generative AI uses natural language processing technology to extract information from the audio and subtitles in the video, supplementing its understanding of the operation procedures. As a result, the analysis unit can analyze the content of the video from multiple angles and accurately detect detailed operation procedures. The analysis unit also saves the analysis results in a database so that subsequent generation and creation units can easily access them. As a result, the analysis unit can analyze the content of the video efficiently and accurately, improving the overall performance of the system.
[0031] The generation unit generates images and explanatory text based on the video content analyzed by the analysis unit. The generation unit generates images and explanatory text for each step, for example, using a generation AI. The generation AI can generate images and explanatory text for each step based on the video content. Specifically, the generation AI selects important frames for each step based on the operating procedures provided by the analysis unit and extracts them as images. Furthermore, the generation AI uses natural language generation technology to generate easy-to-understand explanatory text for each step's operating procedures. For example, the generation AI selects appropriate technical terms and expressions to explain the operating procedures concisely and clearly, creating explanatory text in a user-friendly format. The generation unit also centrally manages the generated images and explanatory text, making them easily accessible to subsequent creation units. This allows the generation unit to efficiently and accurately generate images and explanatory text based on the information provided by the analysis unit, improving the overall system performance.
[0032] The creation unit creates the manual based on the images and explanatory text generated by the generation unit. For example, the creation unit uses a generation AI to create the manual based on the generated images and explanatory text. The generation AI can automatically determine the structure and format of the manual based on the generated images and explanatory text. Specifically, the generation AI arranges the images and explanatory text for each step in a sequential manner, creating a manual that allows users to understand the operating procedures at a glance. Furthermore, the generation AI can optimize the layout and design of the manual according to the user's needs and objectives. For example, the generation AI highlights important information and appropriately places diagrams and icons so that users can quickly understand specific operating procedures. The creation unit also saves the generated manual in digital format, making it easily accessible to users. This allows the creation unit to efficiently and accurately create manuals based on the information provided by the generation unit, improving the overall system performance.
[0033] The support department will establish a chatbot to answer questions about manuals created by the creation department. For example, the support department will use AI to have the chatbot answer questions from users about the contents of the manuals. The AI can provide appropriate answers to user questions. Specifically, the AI uses natural language processing technology to understand user questions and generate appropriate answers. For example, if a user asks about a specific operating procedure, the AI can provide detailed explanations and additional images related to that procedure. The AI can also analyze the user's question history and prepare answers to frequently asked questions in advance, enabling a quick response. Furthermore, the support department can collect user feedback and continuously improve the accuracy and speed of the chatbot's responses. This allows the support department to provide users with quick and appropriate support and improve the overall system performance.
[0034] The analysis unit can analyze the content of a video using a generative AI. For example, the analysis unit can use the generative AI to analyze each frame of a video and detect specific objects or actions. The generative AI can generate images and explanatory text for each step based on the video content. For example, the analysis unit can use the generative AI to detect specific operation steps in a video and generate images and explanatory text corresponding to those steps. This allows for efficient analysis of video content using the generative AI. The generative AI analyzes video content using, for example, a deep learning algorithm. The generative AI can learn from large amounts of video data and provide highly accurate analysis results. For example, the generative AI can detect specific objects or actions in a video with high accuracy and analyze their content. This allows the analysis unit to efficiently analyze video content using the generative AI.
[0035] The generation unit can generate images and descriptions for each step using a generation AI. For example, the generation unit uses a generation AI to generate images and descriptions for each step. The generation AI can generate images and descriptions for each step based on the content of the video. For example, the generation unit uses a generation AI to generate images and descriptions corresponding to specific operation procedures in the video. This allows for efficient generation of images and descriptions for each step using the generation AI. The generation AI analyzes the video content using natural language processing technology, for example, and generates appropriate descriptions. The generation AI can learn from large amounts of text data and generate highly accurate descriptions. For example, the generation AI generates highly accurate descriptions corresponding to specific operation procedures in the video. This allows the generation unit to efficiently generate images and descriptions for each step using the generation AI.
[0036] The creation unit can create manuals based on images and descriptions generated by a generative AI. For example, the creation unit uses the generative AI to create manuals based on the generated images and descriptions. The generative AI can automatically determine the structure and format of the manual based on the generated images and descriptions. For example, the creation unit uses the generative AI to arrange the images and descriptions for each step in a sequential manner, creating a manual where the operating procedures are immediately clear. This allows for efficient manual creation using the generative AI. The generative AI uses, for example, a machine learning algorithm to create manuals based on the generated images and descriptions. The generative AI can learn from large amounts of manual data and create highly accurate manuals. For example, the generative AI arranges the images and descriptions for each step in the optimal order, creating an easy-to-understand manual. This allows the creation unit to efficiently create manuals using the generative AI.
[0037] The support department can use a chatbot to answer user questions about the manual's content. For example, the support department can use AI to have a chatbot respond to user questions about the manual's content. The AI can provide appropriate answers to user questions. For example, the support department can use AI to provide detailed explanations for questions about specific operating procedures. This allows the chatbot to respond to user questions quickly. The AI can analyze user questions and generate appropriate answers using, for example, natural language processing technology. The AI can learn from large amounts of question data and provide highly accurate answers. For example, the AI can quickly provide relevant information in response to user questions. This allows the support department to respond to user questions quickly using AI.
[0038] The support department can use a generation AI to manage versions and always provide the latest manuals. For example, the support department uses a generation AI to manage manual versions. The generation AI can automatically track the manual's change history and manage the latest version. For example, the support department can use the generation AI to automatically collect manual update information and provide the latest manual. This ensures that the latest manuals are always provided using the generation AI. The generation AI manages the manual's change history using a version control system, for example. The generation AI can automatically save each version of the manual and provide the latest version as needed. For example, the generation AI can collect manual update information in real time and provide the latest manual. This ensures that the support department can always provide the latest manuals using the generation AI.
[0039] The analysis unit can detect specific actions and gestures during video analysis and refine the analysis results based on these detections. For example, the analysis unit can use generative AI to detect hand movements in a video and analyze the operation procedures in detail. The generative AI can detect specific actions and gestures based on the content of the video. For example, the analysis unit can use generative AI to detect facial expressions in a video and analyze the user's intent. The analysis unit can also use generative AI to detect body movements in a video and analyze the workflow. This allows for the refinement of analysis results by detecting specific actions and gestures. The generative AI can use computer vision technology, for example, to detect specific actions and gestures in a video with high accuracy. The generative AI can learn from large amounts of video data and perform highly accurate action detection. For example, the generative AI can detect hand movements and facial expressions in a video with high accuracy and analyze their content. This allows the analysis unit to use generative AI to detect specific actions and gestures during video analysis and refine the analysis results based on these detections.
[0040] The analysis unit can analyze background and ambient sounds during video analysis and add information about the work environment. For example, the analysis unit can use a generation AI to analyze background sounds in a video and evaluate the noise level of the work environment. The generation AI can analyze background and ambient sounds based on the content of the video. For example, the analysis unit can use the generation AI to analyze ambient sounds in a video and identify the work location. The analysis unit can also use the generation AI to analyze audio in a video and extract the content of work instructions. In this way, by analyzing background and ambient sounds, information about the work environment can be added. The generation AI can use, for example, speech recognition technology to analyze background and ambient sounds in a video with high accuracy. The generation AI can learn from a large amount of audio data and perform high-precision audio analysis. For example, the generation AI can analyze background and ambient sounds in a video with high accuracy and evaluate their content. In this way, the analysis unit can use the generation AI to analyze background and ambient sounds during video analysis and add information about the work environment.
[0041] The analysis unit can improve analysis accuracy by referring to the user's past video analysis history when analyzing a video. For example, the analysis unit uses a generative AI to refer to the user's past video analysis history. The generative AI can optimize the analysis algorithm based on the past analysis history. For example, the analysis unit uses the generative AI to refer to the user's past analysis history and improve the accuracy when analyzing similar videos. Furthermore, the analysis unit can use the generative AI to optimize the analysis algorithm based on the user's past analysis results. In addition, the analysis unit can use the generative AI to analyze the user's past analysis history and understand analysis trends. This allows for improved analysis accuracy by referring to past analysis history. The generative AI, for example, uses a machine learning algorithm to learn the user's past analysis history. The generative AI can optimize the analysis algorithm based on a large amount of analysis history data. For example, the generative AI analyzes the user's past analysis history with high accuracy and improves analysis accuracy. This allows the analysis unit to improve analysis accuracy by referring to the user's past video analysis history when analyzing a video using the generative AI.
[0042] The analysis unit can customize the analysis results when analyzing videos by taking into account the user's geographical location information. For example, the analysis unit uses a generative AI to consider the user's geographical location information. The generative AI can then customize the analysis results based on the user's location information. For instance, the analysis unit uses the generative AI to analyze region-specific work procedures based on the user's geographical location information. Furthermore, the analysis unit can use the generative AI to refer to the user's location information and provide analysis results tailored to the local language and culture. Additionally, the analysis unit can use the generative AI to perform analysis that considers local environmental conditions based on the user's location information. This allows for the customization of analysis results by considering geographical location information. For example, the generative AI obtains the user's geographical location information using a location information service. The generative AI can then optimize the analysis results based on the location data. For instance, the generative AI obtains the user's location information with high accuracy and customizes the analysis results based on that information. This allows the analysis unit to customize the analysis results when analyzing videos by taking into account the user's geographical location information using the generative AI.
[0043] The generation unit can generate images and explanatory text to highlight specific scenes or important steps in a video during the generation process. For example, the generation unit can use a generation AI to generate images and explanatory text that highlight important operational steps within a video. The generation AI can highlight specific scenes or important steps based on the content of the video. For example, the generation unit can use the generation AI to capture a specific scene in a video and add a detailed explanatory text. The generation unit can also use the generation AI to highlight important points in a video, visually emphasizing them. This allows for the provision of easy-to-understand manuals by highlighting specific scenes or important steps. The generation AI can use computer vision technology, for example, to detect specific scenes or important steps in a video with high accuracy. The generation AI can learn from large amounts of video data and perform high-precision scene detection. For example, the generation AI can detect important operational steps in a video with high accuracy and highlight their content. This allows the generation unit to use the generation AI to generate images and explanatory text to highlight specific scenes or important steps in a video.
[0044] The generation unit can generate a description containing interactive elements based on the video content during the generation process. For example, the generation unit can use a generation AI to generate a description containing interactive elements (e.g., clickable links and buttons) based on the video content. The generation AI can generate a description containing interactive elements based on the video content. For example, the generation unit can use the generation AI to add links to specific operation steps within the video and provide detailed explanations. Furthermore, the generation unit can use the generation AI to place buttons at important scenes within the video and display related information. In addition, the generation unit can use the generation AI to add interactive elements to each step within the video, allowing users to access detailed information. This makes it easier for users to access detailed information by generating a description containing interactive elements. The generation AI can use web technologies, for example, to generate a description containing interactive elements. The generation AI can learn from large amounts of web data and generate highly accurate interactive elements. For example, the generation AI can accurately add links and buttons to specific operation steps within the video and make their content interactive. This allows the generation unit to use generation AI to generate explanatory text that includes interactive elements based on the video content.
[0045] The generation unit can customize the generated content by referring to the user's past manual usage history during generation. For example, the generation unit uses a generation AI to refer to the user's past manual usage history. The generation AI can customize the generated content based on past usage history. For example, the generation unit uses the generation AI to refer to the user's past manual usage history and customize and generate similar content. Furthermore, the generation unit can use the generation AI to generate optimal explanatory text and images based on the user's past usage history. In addition, the generation unit can use the generation AI to analyze the user's past usage history and generate manuals tailored to the user's preferences. This allows for customization of the generated content by referring to past usage history. The generation AI, for example, uses a machine learning algorithm to learn the user's past usage history. The generation AI can optimize the generated content based on a large amount of usage history data. For example, the generation AI analyzes the user's past usage history with high accuracy and customizes the generated content. This allows the generation unit to customize the generated content by referring to the user's past manual usage history during generation using the generation AI.
[0046] The generation unit can generate optimal images and descriptions by considering the user's device information during the generation process. For example, the generation unit uses a generation AI to consider the user's device information. The generation AI can generate optimal images and descriptions based on the user's device information. For example, the generation unit uses the generation AI to generate images and descriptions that match the screen size of the user's device. Furthermore, the generation unit can use the generation AI to consider the performance of the user's device and generate images and descriptions in the optimal format. In addition, the generation unit can use the generation AI to refer to the user's device usage and select the optimal display method. This allows for the generation of optimal images and descriptions by considering device information. The generation AI can, for example, obtain the user's device information using a device information service. The generation AI can optimize the generated content based on the device information data. For example, the generation AI obtains the user's device information with high accuracy and generates optimal images and descriptions based on that information. This allows the generation unit to generate optimal images and descriptions by considering the user's device information during the generation process using the generation AI.
[0047] The creation department can incorporate improvements to the manual by referring to past user feedback when creating it. For example, the creation department can use generative AI to refer to past user feedback. The generative AI can then incorporate improvements to the manual based on that past feedback. For example, the creation department can use generative AI to refer to past user feedback and create a manual that incorporates those improvements. Furthermore, the creation department can use generative AI to optimize the manual's content based on user feedback. In addition, the creation department can use generative AI to analyze past user feedback and create a manual that meets user needs. This allows for improvements to the manual by referring to past feedback. The generative AI learns from past user feedback using machine learning algorithms, for example. Based on a large amount of feedback data, the generative AI can optimize improvements to the manual. For example, the generative AI analyzes past user feedback with high accuracy and incorporates improvements to the manual. This allows the creation department to use generative AI to refer to past user feedback and incorporate improvements to the manual when creating it.
[0048] The creation unit can use customized templates based on the video content when creating manuals. For example, the creation unit uses a generative AI to create customized templates based on the video content. The generative AI can select the optimal template based on the video content. For example, the creation unit can use the generative AI to analyze the video content, select the optimal template, and create the manual. Furthermore, the creation unit can use the generative AI to create a manual using templates corresponding to each step of the video. Additionally, the creation unit can use the generative AI to generate customized templates based on the video content and create the manual. This allows for the provision of more appropriate manuals by using customized templates. The generative AI, for example, uses a template generation algorithm to generate customized templates based on the video content. The generative AI can learn from a large amount of template data and generate highly accurate templates. For example, the generative AI can analyze the video content with high accuracy and generate the optimal template based on that content. This allows the creation unit to use the generative AI to create customized templates based on the video content when creating manuals.
[0049] The creation unit can create optimal manuals by considering the user's geographical location information during manual creation. For example, the creation unit uses a generative AI to consider the user's geographical location information. The generative AI can create optimal manuals based on the user's location information. For example, the creation unit uses the generative AI to create manuals that include region-specific work procedures based on the user's geographical location information. Furthermore, the creation unit can use the generative AI to refer to the user's location information and create manuals tailored to the local language and culture. In addition, the creation unit can use the generative AI to create manuals that consider local environmental conditions based on the user's location information. This allows for the creation of optimal manuals by considering geographical location information. The generative AI, for example, uses a location information service to acquire the user's geographical location information. The generative AI can optimize the manual content based on the location data. For example, the generative AI acquires the user's location information with high accuracy and creates an optimal manual based on that information. This allows the creation unit to create optimal manuals by considering the user's geographical location information during manual creation using the generative AI.
[0050] The creation department can analyze users' social media activity and include relevant information in the manual when creating it. For example, the creation department can use generative AI to analyze users' social media activity. Based on this social media activity, the generative AI can include relevant information in the manual. For example, the creation department can use generative AI to analyze users' social media activity and reflect relevant information in the manual. Furthermore, the creation department can use generative AI to refer to users' social media feedback and optimize the manual's content. In addition, the creation department can use generative AI to create a manual tailored to the user's preferences based on their social media activity. This allows for the inclusion of relevant information in the manual by analyzing social media activity. For example, the generative AI can use social media analysis algorithms to analyze users' social media activity. The generative AI can learn from large amounts of social media data and perform highly accurate analysis. For example, the generative AI can analyze users' social media activity with high accuracy and reflect that information in the manual. This allows the creation department to use generative AI to analyze users' social media activity and include relevant information in the manual when creating it.
[0051] The support department can provide the best possible answer when the chatbot responds by referring to the user's past question history. For example, the support department uses generative AI to refer to the user's past question history. Based on this past question history, the generative AI can provide the best possible answer. For example, the support department can use generative AI to refer to the user's past question history and provide the best possible answer to similar questions. Furthermore, the support department can use generative AI to improve the accuracy of answers based on the user's past question history. In addition, the support department can use generative AI to analyze the user's past question history and provide answers tailored to the user's tendencies. This allows the support department to provide the best possible answer by referring to the past question history. For example, the generative AI can analyze the user's past question history using a question history analysis algorithm. The generative AI can learn from large amounts of question history data and perform highly accurate analysis. For example, the generative AI can analyze the user's past question history with high accuracy and provide the best possible answer based on that information. This allows the support department to use generative AI to provide the best possible answer when the chatbot responds by referring to the user's past question history.
[0052] The support department can provide customized responses to chatbots based on the user's current situation and environment. For example, the support department uses generative AI to analyze the user's current situation and environment. The generative AI can then provide the optimal response based on the user's situation and environment. For instance, the support department can use generative AI to analyze the user's current situation and provide the optimal response. Furthermore, the support department can use generative AI to refer to the user's environmental information and provide responses appropriate to the environment. In addition, the support department can use generative AI to provide customized responses based on the user's current situation. This allows for the provision of more appropriate responses by providing customized responses based on the current situation and environment. For example, the generative AI uses a situation analysis algorithm to analyze the user's current situation. The generative AI can learn from large amounts of situational data and perform highly accurate analysis. For example, the generative AI can analyze the user's current situation with high accuracy and provide the optimal response based on that information. This allows the support department to use generative AI to provide customized responses to chatbots based on the user's current situation and environment when they respond.
[0053] The support department can provide optimal answers when the chatbot responds, taking into account the user's geographical location. For example, the support department uses generative AI to consider the user's geographical location. The generative AI can provide optimal answers based on the user's location. For example, the support department can use generative AI to provide answers that include region-specific information based on the user's geographical location. Furthermore, the support department can use generative AI to refer to the user's location and provide answers tailored to the local language and culture. In addition, the support department can use generative AI to provide answers that consider local environmental conditions based on the user's location. This allows for the provision of optimal answers by considering geographical location. The generative AI can, for example, obtain the user's geographical location using location services. The generative AI can optimize the content of the answers based on the location data. For example, the generative AI can obtain the user's location with high accuracy and provide optimal answers based on that information. This allows the support department to use generative AI to provide optimal answers when the chatbot responds, taking into account the user's geographical location.
[0054] The support department can analyze the user's social media activity and provide relevant information when the chatbot responds. For example, the support department can use generative AI to analyze the user's social media activity. Based on this social media activity, the generative AI can provide relevant information. For example, the support department can use generative AI to analyze the user's social media activity and reflect relevant information in the response. Furthermore, the support department can use generative AI to refer to the user's social media feedback and optimize the content of the response. In addition, the support department can use generative AI to provide responses tailored to the user's preferences based on their social media activity. This allows the support department to provide relevant information by analyzing social media activity. For example, the generative AI can use social media analysis algorithms to analyze the user's social media activity. The generative AI can learn from large amounts of social media data and perform highly accurate analysis. For example, the generative AI can analyze the user's social media activity with high accuracy and provide the optimal response based on that information. This allows the support department to use generative AI to analyze the user's social media activity and provide relevant information when the chatbot responds.
[0055] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0056] The analysis unit can detect specific actions and gestures during video analysis and refine the analysis results based on these detections. For example, it can use generative AI to detect hand movements in a video and analyze the operation procedures in detail. It can also detect facial expressions in a video and analyze the user's intentions. Furthermore, it can detect body movements in a video and analyze the workflow. This allows for the refinement of analysis results by detecting specific actions and gestures.
[0057] The analysis unit can analyze background and ambient sounds during video analysis and add information about the work environment. For example, it can use a generation AI to analyze background sounds in a video and evaluate the noise level of the work environment. It can also analyze ambient sounds in a video to identify the work location. Furthermore, it can analyze the audio in a video and extract the content of work instructions. In this way, by analyzing background and ambient sounds, information about the work environment can be added.
[0058] The analysis unit can improve analysis accuracy by referring to the user's past video analysis history during video analysis. For example, it can use a generation AI to refer to the user's past analysis history and improve the accuracy of analyzing similar videos. It can also optimize the analysis algorithm based on the user's past analysis results. Furthermore, it can analyze the user's past analysis history to understand analysis trends. As a result, analysis accuracy can be improved by referring to past analysis history.
[0059] The generation unit can generate images and explanatory text to highlight specific scenes or important steps in a video during the generation process. For example, it can use generation AI to generate images and explanatory text that highlight important operational steps within a video. It can also capture specific scenes within a video and add detailed explanatory text. Furthermore, it can highlight important points within a video for visual emphasis. This allows for the provision of easy-to-understand manuals by highlighting specific scenes or important steps.
[0060] The generation unit can generate explanatory text that includes interactive elements based on the video content during the generation process. For example, it can use generation AI to generate explanatory text that includes interactive elements (such as clickable links and buttons) based on the video content. It can also add links to specific operation steps within the video to provide detailed explanations. Furthermore, it can place buttons at important scenes within the video to display related information. By generating explanatory text that includes interactive elements, it makes it easier for users to access detailed information.
[0061] The following briefly describes the processing flow for example form 1.
[0062] Step 1: The analysis unit analyzes the video. The analysis unit uses a generation AI to analyze the content of the video, and can analyze each frame of the video to detect specific objects or actions. For example, it can detect a specific operation procedure within the video and generate an image and explanatory text corresponding to that procedure. Step 2: The generation unit generates images and descriptions based on the video content analyzed by the analysis unit. The generation unit uses a generation AI to generate images and descriptions for each step, and generates images and descriptions corresponding to specific operation procedures within the video. Step 3: The creation unit creates the manual based on the images and explanatory text generated by the generation unit. The creation unit automatically determines the structure and format of the manual based on the images and explanatory text generated using the generation AI, and arranges the images and explanatory text for each step in order to create a manual that allows users to understand the operating procedures at a glance. Step 4: The support department will set up a chatbot to answer questions about the manuals created by the creation department. The support department will use AI to have the chatbot answer user questions about the manual's content and provide detailed explanations for questions about specific operating procedures.
[0063] (Example of form 2) The manual creation system according to an embodiment of the present invention is a system that automatically creates a manual based on video footage. This manual creation system uses a generative AI to generate images and text, instantly creating a manual that is easy for anyone to understand. Furthermore, by providing a chatbot that answers questions about the manual, a more detailed understanding can be obtained. The manual creation system according to an embodiment of the present invention is a system that automatically creates a manual based on video footage. This system uses a generative AI to generate images and text, instantly creating a manual that is easy for anyone to understand. Furthermore, by providing a chatbot that answers questions about the manual, a more detailed understanding can be obtained. First, the user shoots a video. This video is input into the generative AI. The generative AI analyzes the content of the video and generates images and explanatory text for each step. For example, if a video of the machine operation procedure is input into the generative AI, the generative AI generates images and explanatory text for each operation step. Next, the manual is automatically created based on the images and explanatory text generated by the generative AI. The generated manual is structured to be easy for first-time users to understand. For example, the images and explanatory text for each step are arranged in order, so that the operation procedure can be understood at a glance. Furthermore, the generated manuals include a chatbot. When a user asks a question about the manual's content, the chatbot provides an answer. For example, if a user asks about a specific operating procedure, the chatbot will provide a detailed explanation of that procedure. This system significantly reduces the effort required to create manuals and simplifies version control. Because the generation AI automatically creates the manuals, the latest version is always provided. Also, because a chatbot is included, users can get answers to any questions they may have immediately. For example, in the manufacturing industry, when creating manuals for the operation of new machinery, it was previously necessary to create the manuals manually. However, by using this service, manuals can be automatically created simply by recording a video. This significantly improves work efficiency, and even first-time users can easily understand the operating procedures.This allows the manual creation system to efficiently provide easy-to-understand manuals by analyzing videos, generating images and explanatory text, creating manuals, and providing support through a chatbot.
[0064] The manual creation system according to this embodiment comprises an analysis unit, a generation unit, a creation unit, and a support unit. The analysis unit analyzes a video. The analysis unit analyzes the content of the video using, for example, a generation AI. The generation AI can analyze each frame of the video and detect specific objects or actions. For example, the analysis unit uses the generation AI to detect a specific operation procedure in the video and generates an image and explanatory text corresponding to that procedure. The generation unit generates images and explanatory text based on the content of the video analyzed by the analysis unit. The generation unit generates images and explanatory text for each step using, for example, a generation AI. The generation AI can generate images and explanatory text for each step based on the content of the video. For example, the generation unit uses the generation AI to generate images and explanatory text corresponding to a specific operation procedure in the video. The creation unit creates a manual based on the images and explanatory text generated by the generation unit. The creation unit creates a manual based on the generated images and explanatory text using, for example, a generation AI. The generation AI can automatically determine the structure and format of the manual based on the generated images and explanatory text. For example, the creation unit uses a generation AI to arrange images and explanatory text for each step in a sequential manner, creating a manual that allows users to understand the operating procedures at a glance. The support unit provides a chatbot that answers questions about the manual created by the creation unit. For example, the support unit uses AI to provide answers when a user asks a question about the contents of the manual. The AI can provide appropriate answers to the user's questions. For example, the support unit uses AI to provide detailed explanations in response to questions about specific operating procedures. Thus, the manual creation system according to this embodiment can efficiently provide easy-to-understand manuals by analyzing videos, generating images and explanatory text, creating manuals, and providing support through a chatbot.
[0065] The analysis unit analyzes the video. For example, the analysis unit uses a generative AI to analyze the content of the video. The generative AI can analyze each frame of the video and detect specific objects or actions. Specifically, the generative AI uses deep learning technology to analyze each frame in the video individually, and while considering the continuity between frames, it detects specific operation procedures and important actions. For example, it recognizes hand movements, tool usage, and changes on the screen in the video with high accuracy and extracts them as operation procedures. Furthermore, the generative AI uses natural language processing technology to extract information from the audio and subtitles in the video, supplementing its understanding of the operation procedures. As a result, the analysis unit can analyze the content of the video from multiple angles and accurately detect detailed operation procedures. The analysis unit also saves the analysis results in a database so that subsequent generation and creation units can easily access them. As a result, the analysis unit can analyze the content of the video efficiently and accurately, improving the overall performance of the system.
[0066] The generation unit generates images and explanatory text based on the video content analyzed by the analysis unit. The generation unit generates images and explanatory text for each step, for example, using a generation AI. The generation AI can generate images and explanatory text for each step based on the video content. Specifically, the generation AI selects important frames for each step based on the operating procedures provided by the analysis unit and extracts them as images. Furthermore, the generation AI uses natural language generation technology to generate easy-to-understand explanatory text for each step's operating procedures. For example, the generation AI selects appropriate technical terms and expressions to explain the operating procedures concisely and clearly, creating explanatory text in a user-friendly format. The generation unit also centrally manages the generated images and explanatory text, making them easily accessible to subsequent creation units. This allows the generation unit to efficiently and accurately generate images and explanatory text based on the information provided by the analysis unit, improving the overall system performance.
[0067] The creation unit creates the manual based on the images and explanatory text generated by the generation unit. For example, the creation unit uses a generation AI to create the manual based on the generated images and explanatory text. The generation AI can automatically determine the structure and format of the manual based on the generated images and explanatory text. Specifically, the generation AI arranges the images and explanatory text for each step in a sequential manner, creating a manual that allows users to understand the operating procedures at a glance. Furthermore, the generation AI can optimize the layout and design of the manual according to the user's needs and objectives. For example, the generation AI highlights important information and appropriately places diagrams and icons so that users can quickly understand specific operating procedures. The creation unit also saves the generated manual in digital format, making it easily accessible to users. This allows the creation unit to efficiently and accurately create manuals based on the information provided by the generation unit, improving the overall system performance.
[0068] The support department will establish a chatbot to answer questions about manuals created by the creation department. For example, the support department will use AI to have the chatbot answer questions from users about the contents of the manuals. The AI can provide appropriate answers to user questions. Specifically, the AI uses natural language processing technology to understand user questions and generate appropriate answers. For example, if a user asks about a specific operating procedure, the AI can provide detailed explanations and additional images related to that procedure. The AI can also analyze the user's question history and prepare answers to frequently asked questions in advance, enabling a quick response. Furthermore, the support department can collect user feedback and continuously improve the accuracy and speed of the chatbot's responses. This allows the support department to provide users with quick and appropriate support and improve the overall system performance.
[0069] The analysis unit can analyze the content of a video using a generative AI. For example, the analysis unit can use the generative AI to analyze each frame of a video and detect specific objects or actions. The generative AI can generate images and explanatory text for each step based on the video content. For example, the analysis unit can use the generative AI to detect specific operation steps in a video and generate images and explanatory text corresponding to those steps. This allows for efficient analysis of video content using the generative AI. The generative AI analyzes video content using, for example, a deep learning algorithm. The generative AI can learn from large amounts of video data and provide highly accurate analysis results. For example, the generative AI can detect specific objects or actions in a video with high accuracy and analyze their content. This allows the analysis unit to efficiently analyze video content using the generative AI.
[0070] The generation unit can generate images and descriptions for each step using a generation AI. For example, the generation unit uses a generation AI to generate images and descriptions for each step. The generation AI can generate images and descriptions for each step based on the content of the video. For example, the generation unit uses a generation AI to generate images and descriptions corresponding to specific operation procedures in the video. This allows for efficient generation of images and descriptions for each step using the generation AI. The generation AI analyzes the video content using natural language processing technology, for example, and generates appropriate descriptions. The generation AI can learn from large amounts of text data and generate highly accurate descriptions. For example, the generation AI generates highly accurate descriptions corresponding to specific operation procedures in the video. This allows the generation unit to efficiently generate images and descriptions for each step using the generation AI.
[0071] The creation unit can create manuals based on images and descriptions generated by a generative AI. For example, the creation unit uses the generative AI to create manuals based on the generated images and descriptions. The generative AI can automatically determine the structure and format of the manual based on the generated images and descriptions. For example, the creation unit uses the generative AI to arrange the images and descriptions for each step in a sequential manner, creating a manual where the operating procedures are immediately clear. This allows for efficient manual creation using the generative AI. The generative AI uses, for example, a machine learning algorithm to create manuals based on the generated images and descriptions. The generative AI can learn from large amounts of manual data and create highly accurate manuals. For example, the generative AI arranges the images and descriptions for each step in the optimal order, creating an easy-to-understand manual. This allows the creation unit to efficiently create manuals using the generative AI.
[0072] The support department can use a chatbot to answer user questions about the manual's content. For example, the support department can use AI to have a chatbot respond to user questions about the manual's content. The AI can provide appropriate answers to user questions. For example, the support department can use AI to provide detailed explanations for questions about specific operating procedures. This allows the chatbot to respond to user questions quickly. The AI can analyze user questions and generate appropriate answers using, for example, natural language processing technology. The AI can learn from large amounts of question data and provide highly accurate answers. For example, the AI can quickly provide relevant information in response to user questions. This allows the support department to respond to user questions quickly using AI.
[0073] The support department can use a generation AI to manage versions and always provide the latest manuals. For example, the support department uses a generation AI to manage manual versions. The generation AI can automatically track the manual's change history and manage the latest version. For example, the support department can use the generation AI to automatically collect manual update information and provide the latest manual. This ensures that the latest manuals are always provided using the generation AI. The generation AI manages the manual's change history using a version control system, for example. The generation AI can automatically save each version of the manual and provide the latest version as needed. For example, the generation AI can collect manual update information in real time and provide the latest manual. This ensures that the support department can always provide the latest manuals using the generation AI.
[0074] The analysis unit can estimate the user's emotions and adjust the video analysis method based on the estimated user emotions. The analysis unit estimates the user's emotions, for example, using generative AI. Generative AI can analyze the user's facial expressions and voice to estimate emotions. For example, if the user is nervous, the analysis unit can use generative AI to analyze the video slowly and provide a detailed explanation. Also, if the user is relaxed, the analysis unit can use generative AI to analyze the video quickly and provide a concise explanation. Furthermore, if the user is in a hurry, the analysis unit can use generative AI to focus the analysis on important points. This allows for more appropriate analysis results by adjusting the video analysis method according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input user facial expression data into a generating AI and have the generating AI perform emotion estimation.
[0075] The analysis unit can detect specific actions and gestures during video analysis and refine the analysis results based on these detections. For example, the analysis unit can use generative AI to detect hand movements in a video and analyze the operation procedures in detail. The generative AI can detect specific actions and gestures based on the content of the video. For example, the analysis unit can use generative AI to detect facial expressions in a video and analyze the user's intent. The analysis unit can also use generative AI to detect body movements in a video and analyze the workflow. This allows for the refinement of analysis results by detecting specific actions and gestures. The generative AI can use computer vision technology, for example, to detect specific actions and gestures in a video with high accuracy. The generative AI can learn from large amounts of video data and perform highly accurate action detection. For example, the generative AI can detect hand movements and facial expressions in a video with high accuracy and analyze their content. This allows the analysis unit to use generative AI to detect specific actions and gestures during video analysis and refine the analysis results based on these detections.
[0076] The analysis unit can analyze background and ambient sounds during video analysis and add information about the work environment. For example, the analysis unit can use a generation AI to analyze background sounds in a video and evaluate the noise level of the work environment. The generation AI can analyze background and ambient sounds based on the content of the video. For example, the analysis unit can use the generation AI to analyze ambient sounds in a video and identify the work location. The analysis unit can also use the generation AI to analyze audio in a video and extract the content of work instructions. In this way, by analyzing background and ambient sounds, information about the work environment can be added. The generation AI can use, for example, speech recognition technology to analyze background and ambient sounds in a video with high accuracy. The generation AI can learn from a large amount of audio data and perform high-precision audio analysis. For example, the generation AI can analyze background and ambient sounds in a video with high accuracy and evaluate their content. In this way, the analysis unit can use the generation AI to analyze background and ambient sounds during video analysis and add information about the work environment.
[0077] The analysis unit can estimate the user's emotions and determine the priority of analysis results based on the estimated emotions. The analysis unit estimates the user's emotions, for example, using generative AI. Generative AI can estimate emotions by analyzing the user's facial expressions and voice. For example, the analysis unit can use generative AI to prioritize displaying important analysis results when the user is feeling stressed. Also, the analysis unit can use generative AI to sequentially display detailed analysis results when the user is relaxed. Furthermore, the analysis unit can use generative AI to quickly display concise analysis results when the user is in a hurry. In this way, by prioritizing analysis results according to the user's emotions, important information can be provided preferentially. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input user facial expression data into a generating AI and have the generating AI perform emotion estimation.
[0078] The analysis unit can improve analysis accuracy by referring to the user's past video analysis history when analyzing a video. For example, the analysis unit uses a generative AI to refer to the user's past video analysis history. The generative AI can optimize the analysis algorithm based on the past analysis history. For example, the analysis unit uses the generative AI to refer to the user's past analysis history and improve the accuracy when analyzing similar videos. Furthermore, the analysis unit can use the generative AI to optimize the analysis algorithm based on the user's past analysis results. In addition, the analysis unit can use the generative AI to analyze the user's past analysis history and understand analysis trends. This allows for improved analysis accuracy by referring to past analysis history. The generative AI, for example, uses a machine learning algorithm to learn the user's past analysis history. The generative AI can optimize the analysis algorithm based on a large amount of analysis history data. For example, the generative AI analyzes the user's past analysis history with high accuracy and improves analysis accuracy. This allows the analysis unit to improve analysis accuracy by referring to the user's past video analysis history when analyzing a video using the generative AI.
[0079] The analysis unit can customize the analysis results when analyzing videos by taking into account the user's geographical location information. For example, the analysis unit uses a generative AI to consider the user's geographical location information. The generative AI can then customize the analysis results based on the user's location information. For instance, the analysis unit uses the generative AI to analyze region-specific work procedures based on the user's geographical location information. Furthermore, the analysis unit can use the generative AI to refer to the user's location information and provide analysis results tailored to the local language and culture. Additionally, the analysis unit can use the generative AI to perform analysis that considers local environmental conditions based on the user's location information. This allows for the customization of analysis results by considering geographical location information. For example, the generative AI obtains the user's geographical location information using a location information service. The generative AI can then optimize the analysis results based on the location data. For instance, the generative AI obtains the user's location information with high accuracy and customizes the analysis results based on that information. This allows the analysis unit to customize the analysis results when analyzing videos by taking into account the user's geographical location information using the generative AI.
[0080] The generation unit can estimate the user's emotions and adjust the presentation of the generated images and descriptions based on the estimated emotions. For example, the generation unit uses a generation AI to estimate the user's emotions. The generation AI can analyze the user's facial expressions and voice to estimate emotions. For example, the generation unit can use the generation AI to generate simple, highly visible images and descriptions when the user is tense. It can also use the generation AI to generate detailed, colorful images and descriptions when the user is relaxed. Furthermore, it can use the generation AI to generate concise, to the point, descriptions and images when the user is in a hurry. This allows for the provision of more appropriate manuals by adjusting the presentation of images and descriptions according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the generation unit may be performed using AI, or without AI. For example, the generation unit can input user facial expression data into the generation AI and have the generation AI perform emotion estimation.
[0081] The generation unit can generate images and explanatory text to highlight specific scenes or important steps in a video during the generation process. For example, the generation unit can use a generation AI to generate images and explanatory text that highlight important operational steps within a video. The generation AI can highlight specific scenes or important steps based on the content of the video. For example, the generation unit can use the generation AI to capture a specific scene in a video and add a detailed explanatory text. The generation unit can also use the generation AI to highlight important points in a video, visually emphasizing them. This allows for the provision of easy-to-understand manuals by highlighting specific scenes or important steps. The generation AI can use computer vision technology, for example, to detect specific scenes or important steps in a video with high accuracy. The generation AI can learn from large amounts of video data and perform high-precision scene detection. For example, the generation AI can detect important operational steps in a video with high accuracy and highlight their content. This allows the generation unit to use the generation AI to generate images and explanatory text to highlight specific scenes or important steps in a video.
[0082] The generation unit can generate a description containing interactive elements based on the video content during the generation process. For example, the generation unit can use a generation AI to generate a description containing interactive elements (e.g., clickable links and buttons) based on the video content. The generation AI can generate a description containing interactive elements based on the video content. For example, the generation unit can use the generation AI to add links to specific operation steps within the video and provide detailed explanations. Furthermore, the generation unit can use the generation AI to place buttons at important scenes within the video and display related information. In addition, the generation unit can use the generation AI to add interactive elements to each step within the video, allowing users to access detailed information. This makes it easier for users to access detailed information by generating a description containing interactive elements. The generation AI can use web technologies, for example, to generate a description containing interactive elements. The generation AI can learn from large amounts of web data and generate highly accurate interactive elements. For example, the generation AI can accurately add links and buttons to specific operation steps within the video and make their content interactive. This allows the generation unit to use generation AI to generate explanatory text that includes interactive elements based on the video content.
[0083] The generation unit can estimate the user's emotions and adjust the length of the generated images and descriptions based on the estimated emotions. The generation unit estimates the user's emotions, for example, using a generation AI. The generation AI can analyze the user's facial expressions and voice to estimate emotions. For example, the generation unit can use the generation AI to generate short, concise descriptions and images when the user is in a hurry. The generation unit can also use the generation AI to generate detailed descriptions and images when the user is relaxed. Furthermore, the generation unit can use the generation AI to generate descriptions and images with visually stimulating effects when the user is excited. This allows for the provision of more appropriate manuals by adjusting the length of images and descriptions according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input user facial expression data into the generation AI and have the generation AI perform emotion estimation.
[0084] The generation unit can customize the generated content by referring to the user's past manual usage history during generation. For example, the generation unit uses a generation AI to refer to the user's past manual usage history. The generation AI can customize the generated content based on past usage history. For example, the generation unit uses the generation AI to refer to the user's past manual usage history and customize and generate similar content. Furthermore, the generation unit can use the generation AI to generate optimal explanatory text and images based on the user's past usage history. In addition, the generation unit can use the generation AI to analyze the user's past usage history and generate manuals tailored to the user's preferences. This allows for customization of the generated content by referring to past usage history. The generation AI, for example, uses a machine learning algorithm to learn the user's past usage history. The generation AI can optimize the generated content based on a large amount of usage history data. For example, the generation AI analyzes the user's past usage history with high accuracy and customizes the generated content. This allows the generation unit to customize the generated content by referring to the user's past manual usage history during generation using the generation AI.
[0085] The generation unit can generate optimal images and descriptions by considering the user's device information during the generation process. For example, the generation unit uses a generation AI to consider the user's device information. The generation AI can generate optimal images and descriptions based on the user's device information. For example, the generation unit uses the generation AI to generate images and descriptions that match the screen size of the user's device. Furthermore, the generation unit can use the generation AI to consider the performance of the user's device and generate images and descriptions in the optimal format. In addition, the generation unit can use the generation AI to refer to the user's device usage and select the optimal display method. This allows for the generation of optimal images and descriptions by considering device information. The generation AI can, for example, obtain the user's device information using a device information service. The generation AI can optimize the generated content based on the device information data. For example, the generation AI obtains the user's device information with high accuracy and generates optimal images and descriptions based on that information. This allows the generation unit to generate optimal images and descriptions by considering the user's device information during the generation process using the generation AI.
[0086] The creation unit can estimate the user's emotions and adjust the manual's structure based on the estimated emotions. The creation unit can estimate the user's emotions, for example, using generative AI. Generative AI can analyze the user's facial expressions and voice to estimate emotions. For example, the creation unit can use generative AI to create a simple and highly visual manual if the user is nervous. It can also use generative AI to create a detailed and colorful manual if the user is relaxed. Furthermore, it can use generative AI to create a concise and to-the-point manual if the user is in a hurry. This allows for the provision of more appropriate manuals by adjusting the manual's structure according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the creation unit may be performed using AI, or not. For example, the creation unit can input user facial expression data into a generation AI and have the generation AI perform emotion estimation.
[0087] The creation department can incorporate improvements to the manual by referring to past user feedback when creating it. For example, the creation department can use generative AI to refer to past user feedback. The generative AI can then incorporate improvements to the manual based on that past feedback. For example, the creation department can use generative AI to refer to past user feedback and create a manual that incorporates those improvements. Furthermore, the creation department can use generative AI to optimize the manual's content based on user feedback. In addition, the creation department can use generative AI to analyze past user feedback and create a manual that meets user needs. This allows for improvements to the manual by referring to past feedback. The generative AI learns from past user feedback using machine learning algorithms, for example. Based on a large amount of feedback data, the generative AI can optimize improvements to the manual. For example, the generative AI analyzes past user feedback with high accuracy and incorporates improvements to the manual. This allows the creation department to use generative AI to refer to past user feedback and incorporate improvements to the manual when creating it.
[0088] The creation unit can use customized templates based on the video content when creating manuals. For example, the creation unit uses a generative AI to create customized templates based on the video content. The generative AI can select the optimal template based on the video content. For example, the creation unit can use the generative AI to analyze the video content, select the optimal template, and create the manual. Furthermore, the creation unit can use the generative AI to create a manual using templates corresponding to each step of the video. Additionally, the creation unit can use the generative AI to generate customized templates based on the video content and create the manual. This allows for the provision of more appropriate manuals by using customized templates. The generative AI, for example, uses a template generation algorithm to generate customized templates based on the video content. The generative AI can learn from a large amount of template data and generate highly accurate templates. For example, the generative AI can analyze the video content with high accuracy and generate the optimal template based on that content. This allows the creation unit to use the generative AI to create customized templates based on the video content when creating manuals.
[0089] The creation unit can estimate the user's emotions and determine the priority of manuals based on the estimated emotions. The creation unit estimates the user's emotions, for example, using generative AI. Generative AI can estimate emotions by analyzing the user's facial expressions and voice. For example, the creation unit can use generative AI to prioritize the creation of important manuals when the user is stressed. Also, the creation unit can use generative AI to sequentially create detailed manuals when the user is relaxed. Furthermore, the creation unit can use generative AI to quickly create concise manuals when the user is in a hurry. In this way, by determining the priority of manuals according to the user's emotions, important manuals can be provided preferentially. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the creation unit may be performed using AI, for example, or without AI. For example, the creation unit can input user facial expression data into a generation AI and have the generation AI perform emotion estimation.
[0090] The creation unit can create optimal manuals by considering the user's geographical location information during manual creation. For example, the creation unit uses a generative AI to consider the user's geographical location information. The generative AI can create optimal manuals based on the user's location information. For example, the creation unit uses the generative AI to create manuals that include region-specific work procedures based on the user's geographical location information. Furthermore, the creation unit can use the generative AI to refer to the user's location information and create manuals tailored to the local language and culture. In addition, the creation unit can use the generative AI to create manuals that consider local environmental conditions based on the user's location information. This allows for the creation of optimal manuals by considering geographical location information. The generative AI, for example, uses a location information service to acquire the user's geographical location information. The generative AI can optimize the manual content based on the location data. For example, the generative AI acquires the user's location information with high accuracy and creates an optimal manual based on that information. This allows the creation unit to create optimal manuals by considering the user's geographical location information during manual creation using the generative AI.
[0091] The creation department can analyze users' social media activity and include relevant information in the manual when creating it. For example, the creation department can use generative AI to analyze users' social media activity. Based on this social media activity, the generative AI can include relevant information in the manual. For example, the creation department can use generative AI to analyze users' social media activity and reflect relevant information in the manual. Furthermore, the creation department can use generative AI to refer to users' social media feedback and optimize the manual's content. In addition, the creation department can use generative AI to create a manual tailored to the user's preferences based on their social media activity. This allows for the inclusion of relevant information in the manual by analyzing social media activity. For example, the generative AI can use social media analysis algorithms to analyze users' social media activity. The generative AI can learn from large amounts of social media data and perform highly accurate analysis. For example, the generative AI can analyze users' social media activity with high accuracy and reflect that information in the manual. This allows the creation department to use generative AI to analyze users' social media activity and include relevant information in the manual when creating it.
[0092] The support unit can estimate the user's emotions and adjust the chatbot's response method based on the estimated emotions. The support unit can estimate the user's emotions, for example, using generative AI. Generative AI can analyze the user's facial expressions and voice to estimate emotions. For example, the support unit can use generative AI to have the chatbot provide a calm response if the user is nervous. Also, the support unit can use generative AI to have the chatbot provide a friendly response if the user is relaxed. Furthermore, the support unit can use generative AI to have the chatbot provide a quick and concise response if the user is in a hurry. This allows for more appropriate responses by adjusting the chatbot's response method according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the support unit may be performed using AI, for example, or without AI. For example, the support unit can input user facial expression data into a generating AI and have the AI perform emotion estimation.
[0093] The support department can provide the best possible answer when the chatbot responds by referring to the user's past question history. For example, the support department uses generative AI to refer to the user's past question history. Based on this past question history, the generative AI can provide the best possible answer. For example, the support department can use generative AI to refer to the user's past question history and provide the best possible answer to similar questions. Furthermore, the support department can use generative AI to improve the accuracy of answers based on the user's past question history. In addition, the support department can use generative AI to analyze the user's past question history and provide answers tailored to the user's tendencies. This allows the support department to provide the best possible answer by referring to the past question history. For example, the generative AI can analyze the user's past question history using a question history analysis algorithm. The generative AI can learn from large amounts of question history data and perform highly accurate analysis. For example, the generative AI can analyze the user's past question history with high accuracy and provide the best possible answer based on that information. This allows the support department to use generative AI to provide the best possible answer when the chatbot responds by referring to the user's past question history.
[0094] The support department can provide customized responses to chatbots based on the user's current situation and environment. For example, the support department uses generative AI to analyze the user's current situation and environment. The generative AI can then provide the optimal response based on the user's situation and environment. For instance, the support department can use generative AI to analyze the user's current situation and provide the optimal response. Furthermore, the support department can use generative AI to refer to the user's environmental information and provide responses appropriate to the environment. In addition, the support department can use generative AI to provide customized responses based on the user's current situation. This allows for the provision of more appropriate responses by providing customized responses based on the current situation and environment. For example, the generative AI uses a situation analysis algorithm to analyze the user's current situation. The generative AI can learn from large amounts of situational data and perform highly accurate analysis. For example, the generative AI can analyze the user's current situation with high accuracy and provide the optimal response based on that information. This allows the support department to use generative AI to provide customized responses to chatbots based on the user's current situation and environment when they respond.
[0095] The support unit can estimate the user's emotions and determine the priority of the chatbot's responses based on the estimated emotions. The support unit can estimate the user's emotions, for example, by using generative AI. Generative AI can estimate emotions by analyzing the user's facial expressions and voice. For example, the support unit can use generative AI to have the chatbot prioritize providing important answers when the user is stressed. Also, the support unit can use generative AI to have the chatbot sequentially provide detailed answers when the user is relaxed. Furthermore, the support unit can use generative AI to have the chatbot quickly provide concise answers when the user is in a hurry. In this way, important answers can be prioritized by determining the priority of responses according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the support unit may be performed using AI, for example, or without AI. For example, the support unit can input user facial expression data into a generating AI and have the AI perform emotion estimation.
[0096] The support department can provide optimal answers when the chatbot responds, taking into account the user's geographical location. For example, the support department uses generative AI to consider the user's geographical location. The generative AI can provide optimal answers based on the user's location. For example, the support department can use generative AI to provide answers that include region-specific information based on the user's geographical location. Furthermore, the support department can use generative AI to refer to the user's location and provide answers tailored to the local language and culture. In addition, the support department can use generative AI to provide answers that consider local environmental conditions based on the user's location. This allows for the provision of optimal answers by considering geographical location. The generative AI can, for example, obtain the user's geographical location using location services. The generative AI can optimize the content of the answers based on the location data. For example, the generative AI can obtain the user's location with high accuracy and provide optimal answers based on that information. This allows the support department to use generative AI to provide optimal answers when the chatbot responds, taking into account the user's geographical location.
[0097] The support department can analyze the user's social media activity and provide relevant information when the chatbot responds. For example, the support department can use generative AI to analyze the user's social media activity. Based on this social media activity, the generative AI can provide relevant information. For example, the support department can use generative AI to analyze the user's social media activity and reflect relevant information in the response. Furthermore, the support department can use generative AI to refer to the user's social media feedback and optimize the content of the response. In addition, the support department can use generative AI to provide responses tailored to the user's preferences based on their social media activity. This allows the support department to provide relevant information by analyzing social media activity. For example, the generative AI can use social media analysis algorithms to analyze the user's social media activity. The generative AI can learn from large amounts of social media data and perform highly accurate analysis. For example, the generative AI can analyze the user's social media activity with high accuracy and provide the optimal response based on that information. This allows the support department to use generative AI to analyze the user's social media activity and provide relevant information when the chatbot responds.
[0098] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0099] The analysis unit can estimate the user's emotions and adjust how the analysis results are displayed based on those emotions. For example, if the user is stressed, the analysis unit will highlight and concisely display important information. If the user is relaxed, it can display detailed information sequentially. Furthermore, if the user is in a hurry, it can quickly display information that gets straight to the point. By adjusting how the analysis results are displayed according to the user's emotions, it becomes possible to provide more appropriate information.
[0100] The generation unit can estimate the user's emotions and adjust the style of the generated images and descriptions based on those emotions. For example, if the user is nervous, it can generate simple, highly visible images and descriptions. If the user is relaxed, it can generate detailed, colorful images and descriptions. Furthermore, if the user is in a hurry, it can generate concise, to-the-point descriptions and images. By adjusting the style of images and descriptions according to the user's emotions, it is possible to provide a more appropriate manual.
[0101] The creation process can estimate the user's emotions and adjust the manual's structure based on those estimates. For example, if the user is nervous, a simple and highly visual manual can be created. If the user is relaxed, a detailed and colorful manual can be created. Furthermore, if the user is in a hurry, a concise manual focusing on the key points can be created. By adjusting the manual's structure according to the user's emotions, a more appropriate manual can be provided.
[0102] The support team can estimate the user's emotions and adjust the chatbot's response based on those estimates. For example, if the user is nervous, the chatbot can provide a calm response. If the user is relaxed, the chatbot can provide a friendly response. Furthermore, if the user is in a hurry, the chatbot can provide a quick and concise response. By adjusting the chatbot's response based on the user's emotions, more appropriate answers can be provided.
[0103] The analysis unit can estimate the user's emotions and prioritize the analysis results based on those emotions. For example, if the user is stressed, important analysis results will be displayed first. If the user is relaxed, detailed analysis results can be displayed sequentially. Furthermore, if the user is in a hurry, concise analysis results can be displayed quickly. In this way, by prioritizing analysis results according to the user's emotions, important information can be provided preferentially.
[0104] The analysis unit can detect specific actions and gestures during video analysis and refine the analysis results based on these detections. For example, it can use generative AI to detect hand movements in a video and analyze the operation procedures in detail. It can also detect facial expressions in a video and analyze the user's intentions. Furthermore, it can detect body movements in a video and analyze the workflow. This allows for the refinement of analysis results by detecting specific actions and gestures.
[0105] The analysis unit can analyze background and ambient sounds during video analysis and add information about the work environment. For example, it can use a generation AI to analyze background sounds in a video and evaluate the noise level of the work environment. It can also analyze ambient sounds in a video to identify the work location. Furthermore, it can analyze the audio in a video and extract the content of work instructions. In this way, by analyzing background and ambient sounds, information about the work environment can be added.
[0106] The analysis unit can improve analysis accuracy by referring to the user's past video analysis history during video analysis. For example, it can use a generation AI to refer to the user's past analysis history and improve the accuracy of analyzing similar videos. It can also optimize the analysis algorithm based on the user's past analysis results. Furthermore, it can analyze the user's past analysis history to understand analysis trends. As a result, analysis accuracy can be improved by referring to past analysis history.
[0107] The generation unit can generate images and explanatory text to highlight specific scenes or important steps in a video during the generation process. For example, it can use generation AI to generate images and explanatory text that highlight important operational steps within a video. It can also capture specific scenes within a video and add detailed explanatory text. Furthermore, it can highlight important points within a video for visual emphasis. This allows for the provision of easy-to-understand manuals by highlighting specific scenes or important steps.
[0108] The generation unit can generate explanatory text that includes interactive elements based on the video content during the generation process. For example, it can use generation AI to generate explanatory text that includes interactive elements (such as clickable links and buttons) based on the video content. It can also add links to specific operation steps within the video to provide detailed explanations. Furthermore, it can place buttons at important scenes within the video to display related information. By generating explanatory text that includes interactive elements, it makes it easier for users to access detailed information.
[0109] The following briefly describes the processing flow for example form 2.
[0110] Step 1: The analysis unit analyzes the video. The analysis unit uses a generation AI to analyze the content of the video, and can analyze each frame of the video to detect specific objects or actions. For example, it can detect a specific operation procedure within the video and generate an image and explanatory text corresponding to that procedure. Step 2: The generation unit generates images and descriptions based on the video content analyzed by the analysis unit. The generation unit uses a generation AI to generate images and descriptions for each step, and generates images and descriptions corresponding to specific operation procedures within the video. Step 3: The creation unit creates the manual based on the images and explanatory text generated by the generation unit. The creation unit automatically determines the structure and format of the manual based on the images and explanatory text generated using the generation AI, and arranges the images and explanatory text for each step in order to create a manual that allows users to understand the operating procedures at a glance. Step 4: The support department will set up a chatbot to answer questions about the manuals created by the creation department. The support department will use AI to have the chatbot answer user questions about the manual's content and provide detailed explanations for questions about specific operating procedures.
[0111] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0112] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0113] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0114] Each of the multiple elements described above, including the analysis unit, generation unit, creation unit, and support unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the analysis unit captures video using the camera 42 of the smart device 14 and analyzes the content of the video using the specific processing unit 290 of the data processing unit 12. The generation unit generates images and explanatory text based on the content analyzed by the specific processing unit 290 of the data processing unit 12. The creation unit creates a manual based on the images and explanatory text generated by the specific processing unit 290 of the data processing unit 12. The support unit answers user questions using a chatbot provided by the control unit 46A of the smart device 14. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0115] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0116] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0117] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0118] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0119] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0120] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0121] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0122] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0123] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0124] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0125] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0126] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0127] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0128] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0129] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0130] Each of the multiple elements described above, including the analysis unit, generation unit, creation unit, and support unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the analysis unit captures video using the camera 42 of the smart glasses 214 and analyzes the content of the video using the specific processing unit 290 of the data processing unit 12. The generation unit generates images and explanatory text based on the content analyzed by the specific processing unit 290 of the data processing unit 12. The creation unit creates a manual based on the images and explanatory text generated by the specific processing unit 290 of the data processing unit 12. The support unit answers user questions using a chatbot provided by the control unit 46A of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0131] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0132] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0133] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0134] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0135] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0136] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0137] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0138] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0139] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0140] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0141] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0142] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0143] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0144] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0145] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0146] Each of the multiple elements described above, including the analysis unit, generation unit, creation unit, and support unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the analysis unit captures video using the camera 42 of the headset terminal 314 and analyzes the content of the video using the specific processing unit 290 of the data processing unit 12. The generation unit generates images and explanatory text based on the content analyzed by the specific processing unit 290 of the data processing unit 12. The creation unit creates a manual based on the images and explanatory text generated by the specific processing unit 290 of the data processing unit 12. The support unit answers user questions using a chatbot provided by the control unit 46A of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0147] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0148] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0149] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0150] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0151] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0152] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0153] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0154] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0155] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0156] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0157] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0158] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0159] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0160] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0161] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0162] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0163] Each of the multiple elements described above, including the analysis unit, generation unit, creation unit, and support unit, is implemented in, for example, at least one of the robot 414 and the data processing unit 12. For example, the analysis unit captures video using the camera 42 of the robot 414 and analyzes the content of the video using the specific processing unit 290 of the data processing unit 12. The generation unit generates images and explanatory text based on the content analyzed by the specific processing unit 290 of the data processing unit 12. The creation unit creates a manual based on the images and explanatory text generated by the specific processing unit 290 of the data processing unit 12. The support unit answers user questions using, for example, a chatbot provided by the control unit 46A of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0164] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0165] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0166] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0167] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0168] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0169] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0170] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0171] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0172] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0173] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0174] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0175] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0176] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0177] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0178] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0179] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0180] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0181] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0182] (Note 1) The analysis unit analyzes the video, A generation unit generates images and explanatory text based on the content of the video analyzed by the analysis unit, A creation unit creates a manual based on the images and explanatory text generated by the generation unit, The system includes a support unit equipped with a chatbot that responds to manuals created by the creation unit. A system characterized by the following features. (Note 2) The aforementioned analysis unit, The AI generates the content of the video. The system described in Appendix 1, characterized by the features described herein. (Note 3) The generating unit is The AI generates images and explanatory text for each step. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned creation unit, Create a manual based on images and descriptions generated by a generative AI. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned support unit is When a user asks a question about the contents of the manual, the chatbot will answer. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned support unit is We use AI-generated versions to ensure that the latest manuals are always available. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned analysis unit, The system estimates the user's emotions and adjusts the video analysis method based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned analysis unit, During video analysis, specific actions and gestures are detected, and the analysis results are refined based on these. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned analysis unit, When analyzing videos, background and ambient sounds are analyzed, and information about the work environment is added. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned analysis unit, It estimates the user's emotions and prioritizes the analysis results based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned analysis unit, When analyzing videos, we improve analysis accuracy by referring to the user's past video analysis history. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned analysis unit, When analyzing videos, the analysis results are customized by taking into account the user's geographical location information. The system described in Appendix 1, characterized by the features described herein. (Note 13) The generating unit is It estimates the user's emotions and adjusts the way images and descriptions are presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The generating unit is During generation, images and descriptive text are generated to highlight specific scenes or important steps in the video. The system described in Appendix 1, characterized by the features described herein. (Note 15) The generating unit is During generation, a description containing interactive elements is generated based on the video content. The system described in Appendix 1, characterized by the features described herein. (Note 16) The generating unit is It estimates the user's emotions and adjusts the length of the generated images and descriptions based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The generating unit is During generation, the generated content is customized by referring to the user's past manual usage history. The system described in Appendix 1, characterized by the features described herein. (Note 18) The generating unit is During generation, the system takes into account the user's device information to create the most suitable image and description. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned creation unit, We estimate the user's emotions and adjust the manual's structure based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned creation unit, When creating manuals, we refer to past user feedback and incorporate improvements to the manuals accordingly. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned creation unit, When creating manuals, use customized templates based on the video content. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned creation unit, The system estimates the user's emotions and prioritizes manual content based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned creation unit, When creating manuals, we take the user's geographical location into consideration to create the most optimal manual. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned creation unit, When creating the manual, analyze users' social media activity and include relevant information in the manual. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned support unit is It estimates the user's emotions and adjusts the chatbot's response method based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned support unit is When the chatbot responds, it refers to the user's past question history to provide the most appropriate answer. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned support unit is When the chatbot responds, it provides customized answers based on the user's current situation and environment. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned support unit is It estimates the user's emotions and prioritizes the chatbot's responses based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned support unit is When the chatbot responds, it takes the user's geographical location into consideration to provide the most appropriate answer. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned support unit is When the chatbot responds, it analyzes the user's social media activity and provides relevant information. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0183] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. The analysis unit analyzes the video, A generation unit generates images and explanatory text based on the content of the video analyzed by the analysis unit, A creation unit creates a manual based on the images and explanatory text generated by the generation unit, The system includes a support unit equipped with a chatbot that responds to manuals created by the creation unit. A system characterized by the following features.
2. The aforementioned analysis unit, The AI generates the content of the video to analyze it. The system according to feature 1.
3. The generating unit is The AI generates images and explanatory text for each step. The system according to feature 1.
4. The aforementioned creation unit, A manual is created based on images and explanatory text generated by a generative AI. The system according to feature 1.
5. The aforementioned support unit is When a user asks a question about the contents of the manual, the chatbot will answer. The system according to feature 1.
6. The aforementioned support unit is AI-generated versions are used for version control, ensuring that the latest manual is always provided. The system according to feature 1.
7. The aforementioned analysis unit, The system estimates the user's emotions and adjusts the video analysis method based on those estimated emotions. The system according to feature 1.
8. The aforementioned analysis unit, During video analysis, specific actions and gestures are detected, and the analysis results are refined based on these. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A