System
By generating user manuals from videos using generative AI and combining this with chatbot support, the problem of time-consuming and difficult-to-understand user manual creation has been solved, achieving efficient and easy-to-understand user manual creation and instant support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2025-10-15
- Publication Date
- 2026-04-21
AI Technical Summary
In the existing technology, the production of operation manuals requires a lot of time and is difficult for beginners to understand, especially for complex mechanical operation procedures.
Generative AI is used to analyze video content, generate images and explanatory text, and is supported by chatbots to automatically create easy-to-understand user manuals.
This significantly reduces the time spent creating operation manuals, improves operational efficiency, ensures the ease of understanding of operation manuals, and provides instant responses through chatbots. The support department is also able to provide the latest version of the operation manual.
Smart Images

Figure CN121905166A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to a system. Background Technology
[0002] Patent Document 1 discloses a personalized chatbot control method executed by at least one processor, the method comprising: receiving user speech; adding the user speech to a prompt containing instructions related to a chatbot role; encoding the prompt; and inputting the encoded prompt into a language model to generate chatbot speech in response to the user speech.
[0003] Patent document 1: Japanese Patent Application Publication No. 2022-180282. Summary of the Invention
[0004] In existing technologies, creating user manuals requires a lot of time and is difficult for beginners to understand, which presents a problem.
[0005] The system involved in this technical solution is designed to automatically generate easy-to-understand operation manuals based on video data.
[0006] The system involved in this technical solution comprises a parsing unit, a generation unit, a production unit, and a support unit. The parsing unit parses the video. The generation unit generates images and explanatory text based on the video content parsed by the parsing unit. The production unit creates an operation manual based on the images and explanatory text generated by the generation unit. The support unit includes a chatbot that responds to the operation manual created by the production unit.
[0007] The system involved in this technical solution can automatically generate easy-to-understand operation manuals based on video data. Attached Figure Description
[0008] Figure 1 This is a conceptual diagram illustrating an example of the configuration of a data processing system according to the first embodiment.
[0009] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0010] Figure 3 This is a conceptual diagram illustrating an example of the data processing system configuration in the second embodiment.
[0011] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0012] Figure 5 This is a conceptual diagram illustrating an example of the data processing system configuration in the third embodiment.
[0013] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.
[0014] Figure 7 This is a conceptual diagram illustrating an example of the data processing system configuration in the fourth embodiment.
[0015] Figure 8 This is a conceptual diagram illustrating an example of the functions of the main parts of the data processing device and robot according to the fourth embodiment.
[0016] Figure 9 It represents an emotion graph that maps multiple emotions.
[0017] Figure 10 It represents an emotion graph that maps multiple emotions.
[0018] Explanation of reference numerals in the attached figures
[0019] Data processing systems 10, 210, 310, and 410
[0020] 12 Data processing device
[0021] 14 Smart devices
[0022] 214 Smart Glasses
[0023] 314 Head-mounted terminal
[0024] 414 Robot. Detailed Implementation
[0025] Hereinafter, an example of an implementation of the system involved in this disclosure will be described with reference to the accompanying drawings.
[0026] First, let's explain the terms used in the following description.
[0027] In the following embodiments, the processor (hereinafter referred to as "processor"), as indicated by the reference numerals, can be a single computing device or a combination of multiple computing devices. Furthermore, a processor can be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit), etc.
[0028] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory for temporary information storage that is used by the processor as working memory.
[0029] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices used to store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disk (e.g., hard disk), or magnetic tape, etc.
[0030] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface including a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0031] In the following implementation, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects more than three items, the same approach as "A and / or B" applies.
[0032] First Implementation Method
[0033] Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0034] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. An example of the data processing device 12 is a server.
[0035] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0036] The smart device 14 includes a computer 36, a receiver 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiver 38, output device 40, and camera 42 are also connected to the bus 52.
[0037] The receiving device 38 includes a touchscreen 38A and a microphone 38B, etc., for receiving user input. The touchscreen 38A receives user input generated by contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input generated by sound by detecting the user's voice. The control unit 46A sends data representing user input received via the touchscreen 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, a specific processing unit 290 (see reference...) Figure 2 Get the data that represents the user input.
[0038] The output device 40 includes a display 40A and a speaker 40B, etc., and presents data to the user by outputting data in a user-perceptible form (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0039] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0040] Figure 2 An example of the main functions of the data processing device 12 and the smart device 14 is shown.
[0041] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.
[0042] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes inferring and predicting the user's emotions, performing various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).
[0043] In the smart device 14, specific processing is performed by the processor 46. A specific processing program 60 is stored in the memory 50. The specific processing program 60 is used in conjunction with the data processing system 10. The processor 46 reads the specific processing program 60 from the memory 50 and executes the read specific processing program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific processing program 60 executed on the RAM 48. Furthermore, the smart device 14 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.
[0044] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing performed by the data processing system 10 of the first embodiment will be described.
[0045] Implementation Method 1
[0046] The user manual creation system disclosed in this invention is a system for automatically creating user manuals based on captured video data. This system utilizes generative AI to generate images and text, enabling the instant creation of user manuals that are easy for anyone to understand. Furthermore, by setting up a chatbot to respond to the user manual, a more detailed understanding can be obtained. The user first captures a video. This video is input into the generative AI. The generative AI analyzes the video content and generates images and explanatory text for each step. For example, when a video of mechanical operation steps is input into the generative AI, the AI generates images and explanatory text for each operation step. Then, based on the images and explanatory text generated by the generative AI, the user manual is automatically created. The generated user manual is structured to be easy for even beginners to understand. For example, the images and explanatory text for each step are arranged sequentially, making the operation steps clear at a glance. Furthermore, a chatbot is integrated into the generated user manual. When users ask questions about the manual's content, the chatbot provides answers. For example, for questions about specific operating steps, the chatbot offers detailed explanations of those steps. This mechanism significantly reduces the time spent creating user manuals and simplifies version management. Because the generative AI automatically generates the manuals, they are always up-to-date. In addition, the chatbot provides immediate answers to any questions. For instance, in manufacturing, when creating manuals for new machines, manuals used to be created manually; however, with this service, manuals can be automatically created simply by recording videos. This greatly improves efficiency, and even beginners can easily understand the operating steps. Therefore, the user manual creation system, by parsing videos, generating images and explanatory text, creating user manuals, and providing chatbot support, efficiently delivers easy-to-understand user manuals.
[0047] The user manual creation system according to this embodiment includes a parsing unit, a generation unit, a production unit, and a support unit. The parsing unit parses video. For example, the parsing unit may utilize generative AI to parse video content. Generative AI can parse each frame of the video and detect specific objects or actions. For example, the parsing unit may use generative AI to detect specific operation steps in the video and generate images and descriptions corresponding to those steps. The generation unit generates images and descriptions based on the video content parsed by the parsing unit. For example, the generation unit may use generative AI to generate images and descriptions for each step. Generative AI can generate images and descriptions for each step based on video content. For example, the generation unit may use generative AI to generate images and descriptions corresponding to specific operation steps in the video. The production unit creates a user manual based on the images and descriptions generated by the generation unit. For example, the production unit may use generative AI to create a user manual based on the generated images and descriptions. The generative AI can automatically determine the structure and format of the user manual based on the generated images and descriptions. For example, the production unit may use generative AI to arrange the images and descriptions for each step in sequence to create a clear and concise user manual. The support department is equipped with a chatbot that responds to user questions about the operation manuals produced by the production department. For example, the support department can utilize AI to answer user questions about the manual's content. The AI can provide appropriate answers to user questions. For instance, the support department can use AI to provide detailed explanations for questions about specific operating procedures. Therefore, the operation manual production system according to this embodiment, by parsing videos, generating images and explanatory text, producing operation manuals, and providing support through a chatbot, can efficiently provide easy-to-understand operation manuals.
[0048] The analysis unit is used to analyze videos. For example, it can utilize generative AI to analyze video content. Generative AI can analyze each frame of a video, detecting specific objects or actions. Specifically, generative AI uses deep learning technology to analyze each frame of the video individually, while considering the continuity between frames, to detect specific operation steps or important actions. For example, it can accurately identify hand movements, tool usage methods, and changes on the screen, extracting them as operation steps. Furthermore, generative AI also uses natural language processing technology to extract information from the video's audio or subtitles to aid in understanding the operation steps. Thus, the analysis unit can analyze video content from multiple perspectives and accurately detect detailed operation steps. In addition, the analysis unit saves the analysis results to a database for easy access by the generation and production units. Therefore, the analysis unit can efficiently and accurately analyze video content, improving the overall system performance.
[0049] The generation unit generates images and explanatory text based on the video content analyzed by the parsing unit. For example, the generation unit can utilize generative AI to generate images and explanatory text for each step. Specifically, the generative AI selects important frames for each step based on the operation steps provided by the parsing unit and extracts them as images. Furthermore, the generative AI uses natural language generation technology to generate easily understandable explanatory text for each step. For example, the generative AI selects appropriate technical terms and expressions to generate explanatory text for the user in a concise and clear form. In addition, the generation unit manages the generated images and explanatory text centrally for easy access by the production unit. Thus, the generation unit can efficiently and accurately generate images and explanatory text based on the information provided by the parsing unit, improving the overall system performance.
[0050] The production department creates operation manuals based on the images and explanatory text generated by the generation department. For example, the production department can utilize generative AI to create operation manuals based on the generated images and explanatory text. Generative AI can automatically determine the structure and format of the operation manual based on the generated images and explanatory text. Specifically, generative AI arranges the images and explanatory text for each step in sequence, creating a clear and concise operation manual. Furthermore, generative AI can optimize the layout and design of the operation manual according to the user's needs and objectives. For example, generative AI will highlight important information and appropriately configure charts and icons, enabling users to quickly understand specific operation steps. In addition, the production department saves the generated operation manuals digitally for easy user access. Thus, the production department can efficiently and accurately create operation manuals based on the information provided by the generation department, improving the overall system performance.
[0051] The support department has a chatbot that responds to user manuals created by the production department. For example, the support department can utilize AI to answer user questions about the manual content. The AI can provide appropriate answers to user questions. Specifically, the AI uses natural language processing technology to understand user questions and generate appropriate responses. For instance, when a user asks about a specific operation step, the AI can provide a detailed explanation or attached images for that step. Furthermore, the AI analyzes the user's question history to prepare answers to common questions in advance, enabling rapid response. Moreover, the support department can collect user feedback to continuously improve the chatbot's accuracy and response speed. Thus, the support department can provide users with fast and appropriate support, improving the overall system performance.
[0052] The analysis unit can analyze video content using generative AI. For example, it can use generative AI to analyze individual frames of a video and detect specific objects or actions. The generative AI can generate images and descriptions for each step based on the video content. For instance, the analysis unit can use generative AI to detect specific operational steps in a video and generate corresponding images and descriptions. Thus, by utilizing generative AI, video content can be analyzed efficiently. Generative AI can also use deep learning algorithms to analyze video content. It can learn from large amounts of video data, providing high-precision analysis results. For example, it can accurately detect specific objects or actions in a video and analyze their content. Therefore, the analysis unit can efficiently analyze video content using generative AI.
[0053] The generation unit can generate images and explanatory text for each step using generative AI. For example, the generation unit can utilize generative AI to generate images and explanatory text for each step. Generative AI can generate images and explanatory text for each step based on video content. For example, the generation unit can utilize generative AI to generate images and explanatory text corresponding to specific operation steps in a video. Therefore, by utilizing generative AI, images and explanatory text for each step can be generated efficiently. For example, generative AI can use natural language processing technology to parse video content and generate appropriate explanatory text. Generative AI can learn from large amounts of text data to generate highly accurate explanatory text. For example, generative AI can generate highly accurate explanatory text corresponding to specific operation steps in a video. Therefore, the generation unit can efficiently generate images and explanatory text for each step using generative AI.
[0054] The production department can create operation manuals based on generated images and instructions using generative AI. For example, the production department can utilize generative AI to create operation manuals based on generated images and instructions. Generative AI can automatically determine the structure and format of the operation manual based on the generated images and instructions. For instance, the production department can use generative AI to arrange the images and instructions for each step in the optimal order, creating a clear and concise operation manual. Therefore, by utilizing generative AI, operation manuals can be created efficiently. Generative AI can also utilize machine learning algorithms to create operation manuals based on generated images and instructions. Generative AI can learn from a large amount of operation manual data to create highly accurate operation manuals. For instance, generative AI can arrange the images and instructions for each step in the optimal order, creating an easy-to-understand operation manual. Therefore, the production department can create operation manuals efficiently using generative AI.
[0055] The support department can use a chatbot to answer user questions about the user manual. For example, the support department can utilize AI to provide appropriate answers to user questions. For instance, the support department can use AI to provide detailed explanations for questions about specific operating procedures. Thus, by using a chatbot, user questions can be answered quickly. AI can also use natural language processing technology to analyze user questions and generate appropriate responses. AI can learn from large amounts of question data to provide highly accurate answers. For example, AI can quickly provide information relevant to the user's question. Therefore, the support department can quickly answer user questions using AI.
[0056] The support department can use generative AI for version management, ensuring that the latest user manuals are always provided. For example, the support department can utilize generative AI for user manual version management. Generative AI can automatically track the change history of the user manual and manage the latest versions. For instance, the support department can use generative AI to automatically collect update information for the user manual and provide the latest version. Therefore, by utilizing generative AI, the latest user manual can always be provided. Generative AI can also utilize a version management system to manage the change history of the user manual. Generative AI can automatically save various versions of the user manual and provide the latest version as needed. For instance, generative AI can collect update information for the user manual in real time and provide the latest version. Therefore, the support department can always provide the latest user manual through generative AI.
[0057] The analysis unit can detect specific actions or gestures while analyzing videos and refine the analysis results accordingly. For example, the analysis unit can use generative AI to detect hand movements in videos and analyze operation steps in detail. Generative AI can detect specific actions or gestures based on video content. For example, the analysis unit can use generative AI to detect facial expressions in videos and analyze user intentions. Furthermore, the analysis unit can use generative AI to detect body movements in videos and analyze work processes. Thus, by detecting specific actions or gestures, the analysis results can be refined. Generative AI can, for example, use computer vision technology to detect specific actions or gestures in videos with high precision. Generative AI can learn from large amounts of video data to achieve high-precision action detection. For example, generative AI can detect hand movements or facial expressions in videos with high precision and analyze their content. Therefore, the analysis unit can use generative AI to detect specific actions or gestures while analyzing videos and refine the analysis results accordingly.
[0058] The analysis unit can analyze background and ambient sounds while analyzing video, and add information about the work environment. For example, the analysis unit can use generative AI to analyze background sounds in the video and assess the noise level of the work environment. Generative AI can analyze background and ambient sounds based on video content. For example, the analysis unit can use generative AI to analyze ambient sounds in the video to determine the work location. Furthermore, the analysis unit can use generative AI to analyze speech in the video and extract work instructions. Thus, by analyzing background and ambient sounds, information about the work environment can be added. For example, generative AI can use speech recognition technology to analyze background and ambient sounds in the video with high accuracy. Generative AI can learn from large amounts of speech data to achieve high-precision speech analysis. For example, generative AI can analyze background and ambient sounds in the video with high accuracy and evaluate their content. Therefore, the analysis unit can use generative AI to analyze background and ambient sounds while analyzing video and add information about the work environment.
[0059] The parsing department can improve parsing accuracy by referencing the user's past video parsing history. For example, the parsing department can utilize generative AI to reference the user's past video parsing history. Generative AI can optimize the parsing algorithm based on past parsing history. For instance, the parsing department can use generative AI to refer to the user's past parsing history to improve the accuracy when parsing similar videos. Furthermore, the parsing department can use generative AI to optimize the parsing algorithm based on the user's past parsing results. Further, the parsing department can use generative AI to analyze the user's past parsing history to grasp parsing trends. Thus, by referring to past parsing history, parsing accuracy can be improved. For example, generative AI can use machine learning algorithms to learn from the user's past parsing history. Generative AI can optimize the parsing algorithm based on a large amount of parsing history data. For example, generative AI can analyze the user's past parsing history with high precision, improving parsing accuracy. Therefore, the parsing department can improve parsing accuracy by referencing the user's past video parsing history using generative AI.
[0060] The analysis unit can consider the user's geographic location information to customize the analysis results when analyzing video. For example, the analysis unit can utilize generative AI to consider the user's geographic location information. Generative AI can customize analysis results based on the user's location information. For example, the analysis unit can use generative AI to analyze regionally specific work procedures based on the user's geographic location information. Furthermore, the analysis unit can use generative AI to refer to the user's location information to provide analysis results that conform to the regional language and culture. Further, the analysis unit can use generative AI to perform analysis considering regional environmental conditions based on the user's location information. Thus, by considering geographic location information, the analysis results can be customized. For example, generative AI can utilize location information services to obtain the user's geographic location information. Generative AI can optimize the analysis results based on location information data. For example, generative AI can obtain the user's location information with high accuracy and customize the analysis results based on this information. Therefore, the analysis unit can use generative AI to consider the user's geographic location information to customize the analysis results when analyzing video.
[0061] The generation unit can generate images and explanatory text to emphasize specific scenes or important steps in a video during the generation process. For example, the generation unit can utilize generative AI to generate images and explanatory text emphasizing important operational steps in the video. Generative AI can emphasize specific scenes or important steps based on the video content. For example, the generation unit can use generative AI to capture specific scenes in the video and attach detailed explanatory text. Furthermore, the generation unit can use generative AI to highlight important points in the video for visual emphasis. Thus, by emphasizing specific scenes or important steps, an easy-to-understand operation manual can be provided. For example, generative AI can utilize computer vision technology to detect specific scenes or important steps in the video with high precision. Generative AI can learn from large amounts of video data to achieve high-precision scene detection. For example, generative AI can detect important operational steps in a video with high precision and emphasize their content. Therefore, the generation unit can use generative AI to generate images and explanatory text to emphasize specific scenes or important steps in the video during the generation process.
[0062] The generation department can generate descriptive text containing interactive elements based on video content during the generation process. For example, the generation department can utilize generative AI to generate descriptive text containing interactive elements (such as clickable links or buttons) based on video content. Generative AI can generate descriptive text containing interactive elements based on video content. For example, the generation department can use generative AI to add links to specific operation steps in the video, providing detailed explanations. Furthermore, the generation department can use generative AI to set buttons in important scenes in the video, displaying relevant information. Further, the generation department can use generative AI to add interactive elements to each step in the video, allowing users to access detailed information. Thus, by generating descriptive text containing interactive elements, users can more easily access detailed information. Generative AI can, for example, utilize web technologies to generate descriptive text containing interactive elements. Generative AI can learn from large amounts of web data to generate highly accurate interactive elements. For example, generative AI can add links or buttons to specific operation steps in the video with high accuracy, realizing content interactivity. Thus, the generation department can generate descriptive text containing interactive elements based on video content using generative AI.
[0063] The content generation department can customize generated content by referencing the user's past usage history of the user manual during the generation process. For example, the generation department can utilize generative AI to reference the user's past usage history of the user manual. Generative AI can customize generated content based on past usage history. For example, the generation department can use generative AI to reference the user's past usage history of the user manual and customize similar content. Furthermore, the generation department can use generative AI to generate optimal explanatory text and images based on the user's past usage history. Further, the generation department can use generative AI to analyze the user's past usage history and generate user manuals that match user preferences. Thus, by referencing past usage history, the generated content can be customized. For example, generative AI can use machine learning algorithms to learn the user's past usage history. Generative AI can optimize generated content based on a large amount of historical usage data. For example, generative AI can analyze the user's past usage history with high precision and customize generated content. Therefore, the generation department can use generative AI to customize generated content by referencing the user's past usage history of the user manual during the generation process.
[0064] The generation unit can consider the user's device information during generation to generate optimal images and descriptions. For example, the generation unit can utilize generative AI to consider the user's device information. Generative AI can generate optimal images and descriptions based on the user's device information. For example, the generation unit can use generative AI to generate images and descriptions suitable for the user's device screen size. Furthermore, the generation unit can use generative AI to consider the performance of the user's device and generate images and descriptions in the optimal format. Further, the generation unit can use generative AI to refer to the user's device usage and select the optimal display method. Thus, by considering device information, optimal images and descriptions can be generated. For example, generative AI can utilize device information services to obtain the user's device information. Generative AI can optimize the generated content based on device information data. For example, generative AI can obtain the user's device information with high accuracy and generate optimal images and descriptions based on that information. Thus, the generation unit can use generative AI to consider the user's device information during generation to generate optimal images and descriptions.
[0065] The production department can reference past user feedback when creating user manuals to identify areas for improvement. For example, the production department can utilize generative AI to reference past user feedback. Generative AI can then use this feedback to identify areas for improvement in the user manual. For instance, the production department can use generative AI to reference past user feedback and create user manuals that reflect these improvements. Furthermore, the production department can use generative AI to optimize the content of the user manual based on user feedback. Moreover, the production department can use generative AI to analyze past user feedback and create user manuals that meet user needs. Thus, by referencing past feedback, the production department can identify areas for improvement in the user manual. For example, generative AI can use machine learning algorithms to learn from past user feedback. Generative AI can optimize user manuals based on a large amount of feedback data. For instance, generative AI can analyze past user feedback with high precision to identify areas for improvement in the user manual. Therefore, the production department can use generative AI to reference past user feedback when creating user manuals to identify areas for improvement.
[0066] The production department can use templates customized based on video content when creating operation manuals. For example, the production department can utilize generative AI to use templates customized based on video content. Generative AI can select the optimal template based on video content. For example, the production department can use generative AI to analyze video content and select the optimal template to create the operation manual. Furthermore, the production department can use generative AI to create operation manuals using corresponding templates for each step of the video. Further, the production department can use generative AI to generate customized templates based on video content to create operation manuals. Thus, by using customized templates, more suitable operation manuals can be provided. For example, generative AI can use template generation algorithms to generate customized templates based on video content. Generative AI can learn from a large amount of template data to generate high-precision templates. For example, generative AI can analyze video content with high precision and generate the optimal template based on that content. Therefore, the production department can use templates customized based on video content when creating operation manuals using generative AI.
[0067] The production department can consider the user's geographical location information when creating user manuals to produce the optimal manual. For example, the production department can utilize generative AI to consider the user's geographical location information. Generative AI can create the optimal user manual based on the user's location information. For instance, the production department can use generative AI to create a user manual that includes region-specific operating procedures based on the user's geographical location information. Furthermore, the production department can use generative AI to create a user manual that conforms to the region's language and culture, referencing the user's location information. Further, the production department can use generative AI to create a user manual that considers the region's environmental conditions based on the user's location information. Thus, by considering geographical location information, the optimal user manual can be created. For example, generative AI can utilize location information services to obtain the user's geographical location information. Generative AI can optimize the content of the user manual based on location information data. For example, generative AI can obtain the user's location information with high accuracy and create the optimal user manual based on this information. Therefore, the production department can use generative AI to consider the user's geographical location information when creating user manuals to produce the optimal manual.
[0068] The production department can analyze users' social media activities when creating user manuals and incorporate relevant information into the manuals. For example, the production department can utilize generative AI to analyze users' social media activities. Generative AI can incorporate relevant information based on social media activities into the user manual. For example, the production department can use generative AI to analyze users' social media activities and reflect this information in the user manual. Furthermore, the production department can use generative AI to optimize the user manual content by referencing user feedback on social media. Further, the production department can use generative AI to create user manuals that match user preferences based on users' social media activities. Thus, by analyzing social media activities, relevant information can be incorporated into the user manual. For example, generative AI can utilize social media analysis algorithms to analyze users' social media activities. Generative AI can learn from large amounts of social media data to achieve high-precision analysis. For example, generative AI can analyze users' social media activities with high precision and reflect this information in the user manual. Therefore, the production department can use generative AI to analyze users' social media activities when creating user manuals and incorporate relevant information into the manuals.
[0069] The support department can refer to a user's past question history to provide the optimal response when the chatbot answers. For example, the support department can utilize generative AI to reference the user's past question history. Generative AI can provide the optimal response based on past question history. For example, the support department can use generative AI to refer to a user's past question history to provide the optimal response to similar questions. Furthermore, the support department can use generative AI based on the user's past question history to improve the accuracy of the response. Further, the support department can use generative AI to analyze the user's past question history to provide responses that match the user's preferences. Thus, by referring to past question history, the optimal response can be provided. For example, generative AI can use question history analysis algorithms to analyze the user's past question history. Generative AI can learn from a large amount of question history data to achieve high-precision analysis. For example, generative AI can analyze a user's past question history with high precision and provide the optimal response based on this information. Therefore, the support department can use generative AI to refer to the user's past question history to provide the optimal response when the chatbot answers.
[0070] The support department can provide customized responses based on the user's current situation or environment when the chatbot replies. For example, the support department can utilize generative AI to analyze the user's current situation or environment. Generative AI can provide the optimal response based on the user's situation or environment. For example, the support department can use generative AI to analyze the user's current situation and provide the optimal response. Furthermore, the support department can use generative AI to refer to the user's environmental information and provide responses suitable for the environment. Further, the support department can use generative AI to provide customized responses based on the user's current situation. Thus, by providing customized responses based on the current situation or environment, more appropriate responses can be provided. For example, generative AI can use situation analysis algorithms to analyze the user's current situation. Generative AI can learn from large amounts of situation data to achieve high-precision analysis. For example, generative AI can accurately analyze the user's current situation and provide the optimal response based on this information. Therefore, the support department can use generative AI to provide customized responses based on the user's current situation or environment when the chatbot replies.
[0071] The support department can consider the user's geographic location information to provide the optimal response when the chatbot replies. For example, the support department can utilize generative AI to consider the user's geographic location information. Generative AI can provide the optimal response based on the user's location information. For example, the support department can use generative AI to provide responses containing region-specific information based on the user's geographic location information. Furthermore, the support department can use generative AI to provide responses that conform to the region's language and culture, taking into account the user's location information. Further, the support department can use generative AI to provide responses that consider the region's environmental conditions based on the user's location information. Thus, by considering geographic location information, the optimal response can be provided. For example, generative AI can utilize location information services to obtain the user's geographic location information. Generative AI can optimize the response content based on location information data. For example, generative AI can obtain the user's location information with high accuracy and provide the optimal response based on that information. Therefore, the support department can use generative AI to consider the user's geographic location information to provide the optimal response when the chatbot replies.
[0072] The support department can analyze users' social media activity and provide relevant information when the chatbot responds. For example, the support department can utilize generative AI to analyze users' social media activity. Generative AI can provide relevant information based on social media activity. For instance, the support department can use generative AI to analyze users' social media activity and reflect this information in the response. Furthermore, the support department can use generative AI to optimize the response content by referencing user feedback on social media. Further, the support department can use generative AI to provide responses that match user preferences based on users' social media activity. Thus, by analyzing social media activity, relevant information can be provided. For example, generative AI can utilize social media analysis algorithms to analyze users' social media activity. Generative AI can learn from large amounts of social media data to achieve high-precision analysis. For example, generative AI can analyze users' social media activity with high precision and provide the optimal response based on this information. Therefore, the support department can use generative AI to analyze users' social media activity and provide relevant information when the chatbot responds.
[0073] The system involved in this embodiment is not limited to the above examples. For example, various modifications can be made as follows.
[0074] The analysis unit can detect specific actions or gestures while analyzing videos and refine the analysis results accordingly. For example, generative AI can be used to detect hand movements in videos and analyze operation steps in detail. Furthermore, facial expressions in videos can be detected to analyze user intentions. Even further, body movements in videos can be detected to analyze work processes. Thus, by detecting specific actions or gestures, the analysis results can be refined.
[0075] The analysis unit can analyze background and ambient sounds while analyzing video, and add information about the work environment. For example, generative AI can be used to analyze background sounds in the video to assess the noise level of the work environment. Furthermore, ambient sounds in the video can be analyzed to determine the work location. Further, speech in the video can be analyzed to extract work instructions. Thus, by analyzing background and ambient sounds, information about the work environment can be added.
[0076] The parsing department can improve parsing accuracy by referencing a user's past video parsing history. For example, generative AI can be used to improve the accuracy of parsing similar videos by referencing the user's past parsing history. Furthermore, the parsing algorithm can be optimized based on the user's past parsing results. Moreover, the parsing trend can be identified by analyzing the user's past parsing history. Therefore, by referring to past parsing history, parsing accuracy can be improved.
[0077] The generation unit can generate images and explanatory text to emphasize specific scenes or important steps in a video during the generation process. For example, generative AI can be used to generate images and explanatory text that highlight important operational steps in a video. Furthermore, specific scenes in the video can be captured and detailed explanations can be added. Additionally, important points in the video can be highlighted for visual emphasis. Thus, by emphasizing specific scenes or important steps, an easy-to-understand instruction manual can be provided.
[0078] The generation department can generate descriptive text containing interactive elements based on video content during the generation process. For example, generative AI can be used to generate descriptive text with interactive elements (such as clickable links or buttons) based on video content. Furthermore, links can be added to specific operation steps in the video, providing detailed explanations. Additionally, buttons can be placed in important scenes within the video to display relevant information. Thus, by generating descriptive text with interactive elements, users can more easily access detailed information.
[0079] The following is a brief description of the processing flow of Implementation Method 1.
[0080] Step 1: The analysis unit analyzes the video. The analysis unit uses generative AI to analyze the video content, examining each frame and detecting specific objects or actions. For example, it detects specific operational steps in the video and generates corresponding images and descriptions.
[0081] Step 2: The generation unit generates images and explanatory text based on the video content parsed by the parsing unit. The generation unit uses generative AI to generate images and explanatory text for each step, generating images and explanatory text corresponding to specific operation steps in the video.
[0082] Step 3: The production department creates an operation manual based on the images and explanatory text generated by the generation department. The production department uses generative AI to automatically determine the structure and format of the operation manual based on the generated images and explanatory text, arranging the images and explanatory text for each step in sequence to create a clear and concise operation manual.
[0083] Step 4: The support department has set up a chatbot to answer questions in the operation manuals produced by the production department. Utilizing AI, the chatbot responds to user questions about the manual's content, providing detailed explanations for specific operational steps.
[0084] Implementation Method 2
[0085] The user manual creation system disclosed in this invention is a system for automatically creating user manuals based on captured video data. This system utilizes generative AI to generate images and text, enabling the instant creation of user manuals that are easy for anyone to understand. Furthermore, by setting up a chatbot to respond to the user manual, a more detailed understanding can be obtained. The user first captures a video. This video is input into the generative AI. The generative AI analyzes the video content and generates images and explanatory text for each step. For example, when a video of mechanical operation steps is input into the generative AI, the AI generates images and explanatory text for each operation step. Then, based on the images and explanatory text generated by the generative AI, the user manual is automatically created. The generated user manual is structured to be easy for even beginners to understand. For example, the images and explanatory text for each step are arranged sequentially, making the operation steps clear at a glance. Furthermore, a chatbot is integrated into the generated user manual. When users ask questions about the manual's content, the chatbot provides answers. For example, for questions about specific operating steps, the chatbot offers detailed explanations of those steps. This mechanism significantly reduces the time spent creating user manuals and simplifies version management. Because the generative AI automatically generates the manuals, they are always up-to-date. In addition, the chatbot provides immediate answers to any questions. For instance, in manufacturing, when creating manuals for new machines, manuals used to be created manually; however, with this service, manuals can be automatically created simply by recording videos. This greatly improves efficiency, and even beginners can easily understand the operating steps. Therefore, the user manual creation system, by parsing videos, generating images and explanatory text, creating user manuals, and providing chatbot support, efficiently delivers easy-to-understand user manuals.
[0086] The user manual creation system according to this embodiment includes a parsing unit, a generation unit, a production unit, and a support unit. The parsing unit parses video. For example, the parsing unit may utilize generative AI to parse video content. Generative AI can parse each frame of the video and detect specific objects or actions. For example, the parsing unit may use generative AI to detect specific operation steps in the video and generate images and descriptions corresponding to those steps. The generation unit generates images and descriptions based on the video content parsed by the parsing unit. For example, the generation unit may use generative AI to generate images and descriptions for each step. Generative AI can generate images and descriptions for each step based on video content. For example, the generation unit may use generative AI to generate images and descriptions corresponding to specific operation steps in the video. The production unit creates a user manual based on the images and descriptions generated by the generation unit. For example, the production unit may use generative AI to create a user manual based on the generated images and descriptions. The generative AI can automatically determine the structure and format of the user manual based on the generated images and descriptions. For example, the production unit may use generative AI to arrange the images and descriptions for each step in sequence to create a clear and concise user manual. The support department is equipped with a chatbot that responds to user questions about the operation manuals produced by the production department. For example, the support department can utilize AI to answer user questions about the manual's content. The AI can provide appropriate answers to user questions. For instance, the support department can use AI to provide detailed explanations for questions about specific operating procedures. Therefore, the operation manual production system according to this embodiment, by parsing videos, generating images and explanatory text, producing operation manuals, and providing support through a chatbot, can efficiently provide easy-to-understand operation manuals.
[0087] The analysis unit is used to analyze videos. For example, it can utilize generative AI to analyze video content. Generative AI can analyze each frame of a video, detecting specific objects or actions. Specifically, generative AI uses deep learning technology to analyze each frame of the video individually, while considering the continuity between frames, to detect specific operation steps or important actions. For example, it can accurately identify hand movements, tool usage methods, and changes on the screen, extracting them as operation steps. Furthermore, generative AI also uses natural language processing technology to extract information from the video's audio or subtitles to aid in understanding the operation steps. Thus, the analysis unit can analyze video content from multiple perspectives and accurately detect detailed operation steps. In addition, the analysis unit saves the analysis results to a database for easy access by the generation and production units. Therefore, the analysis unit can efficiently and accurately analyze video content, improving the overall system performance.
[0088] The generation unit generates images and explanatory text based on the video content analyzed by the parsing unit. For example, the generation unit can utilize generative AI to generate images and explanatory text for each step. Specifically, the generative AI selects important frames for each step based on the operation steps provided by the parsing unit and extracts them as images. Furthermore, the generative AI uses natural language generation technology to generate easily understandable explanatory text for each step. For example, the generative AI selects appropriate technical terms and expressions to generate explanatory text for the user in a concise and clear form. In addition, the generation unit manages the generated images and explanatory text centrally for easy access by the production unit. Thus, the generation unit can efficiently and accurately generate images and explanatory text based on the information provided by the parsing unit, improving the overall system performance.
[0089] The production department creates operation manuals based on the images and explanatory text generated by the generation department. For example, the production department can utilize generative AI to create operation manuals based on the generated images and explanatory text. Generative AI can automatically determine the structure and format of the operation manual based on the generated images and explanatory text. Specifically, generative AI arranges the images and explanatory text for each step in sequence, creating a clear and concise operation manual. Furthermore, generative AI can optimize the layout and design of the operation manual according to the user's needs and objectives. For example, generative AI will highlight important information and appropriately configure charts and icons, enabling users to quickly understand specific operation steps. In addition, the production department saves the generated operation manuals digitally for easy user access. Thus, the production department can efficiently and accurately create operation manuals based on the information provided by the generation department, improving the overall system performance.
[0090] The support department has a chatbot that responds to user manuals created by the production department. For example, the support department can utilize AI to answer user questions about the manual content. The AI can provide appropriate answers to user questions. Specifically, the AI uses natural language processing technology to understand user questions and generate appropriate responses. For instance, when a user asks about a specific operation step, the AI can provide a detailed explanation or attached images for that step. Furthermore, the AI analyzes the user's question history to prepare answers to common questions in advance, enabling rapid response. Moreover, the support department can collect user feedback to continuously improve the chatbot's accuracy and response speed. Thus, the support department can provide users with fast and appropriate support, improving the overall system performance.
[0091] The analysis unit can analyze video content using generative AI. For example, it can use generative AI to analyze individual frames of a video and detect specific objects or actions. The generative AI can generate images and descriptions for each step based on the video content. For instance, the analysis unit can use generative AI to detect specific operational steps in a video and generate corresponding images and descriptions. Thus, by utilizing generative AI, video content can be analyzed efficiently. Generative AI can also use deep learning algorithms to analyze video content. It can learn from large amounts of video data, providing high-precision analysis results. For example, it can accurately detect specific objects or actions in a video and analyze their content. Therefore, the analysis unit can efficiently analyze video content using generative AI.
[0092] The generation unit can generate images and explanatory text for each step using generative AI. For example, the generation unit can utilize generative AI to generate images and explanatory text for each step. Generative AI can generate images and explanatory text for each step based on video content. For example, the generation unit can utilize generative AI to generate images and explanatory text corresponding to specific operation steps in a video. Therefore, by utilizing generative AI, images and explanatory text for each step can be generated efficiently. For example, generative AI can use natural language processing technology to parse video content and generate appropriate explanatory text. Generative AI can learn from large amounts of text data to generate highly accurate explanatory text. For example, generative AI can generate highly accurate explanatory text corresponding to specific operation steps in a video. Therefore, the generation unit can efficiently generate images and explanatory text for each step using generative AI.
[0093] The production department can create operation manuals based on generated images and instructions using generative AI. For example, the production department can utilize generative AI to create operation manuals based on generated images and instructions. Generative AI can automatically determine the structure and format of the operation manual based on the generated images and instructions. For instance, the production department can use generative AI to arrange the images and instructions for each step in the optimal order, creating a clear and concise operation manual. Therefore, by utilizing generative AI, operation manuals can be created efficiently. Generative AI can also utilize machine learning algorithms to create operation manuals based on generated images and instructions. Generative AI can learn from a large amount of operation manual data to create highly accurate operation manuals. For instance, generative AI can arrange the images and instructions for each step in the optimal order, creating an easy-to-understand operation manual. Therefore, the production department can create operation manuals efficiently using generative AI.
[0094] The support department can use a chatbot to answer user questions about the user manual. For example, the support department can utilize AI to provide appropriate answers to user questions. For instance, the support department can use AI to provide detailed explanations for questions about specific operating procedures. Thus, by using a chatbot, user questions can be answered quickly. AI can also use natural language processing technology to analyze user questions and generate appropriate responses. AI can learn from large amounts of question data to provide highly accurate answers. For example, AI can quickly provide information relevant to the user's question. Therefore, the support department can quickly answer user questions using AI.
[0095] The support department can use generative AI for version management, ensuring that the latest user manuals are always provided. For example, the support department can utilize generative AI for user manual version management. Generative AI can automatically track the change history of the user manual and manage the latest versions. For instance, the support department can use generative AI to automatically collect update information for the user manual and provide the latest version. Therefore, by utilizing generative AI, the latest user manual can always be provided. Generative AI can also utilize a version management system to manage the change history of the user manual. Generative AI can automatically save various versions of the user manual and provide the latest version as needed. For instance, generative AI can collect update information for the user manual in real time and provide the latest version. Therefore, the support department can always provide the latest user manual through generative AI.
[0096] The analysis unit can infer the user's emotions and adjust the video analysis method accordingly. For example, the analysis unit can utilize generative AI to infer the user's emotions. Generative AI can analyze the user's facial expressions and speech to infer emotions. For instance, the analysis unit can use generative AI to slowly analyze the video and provide detailed explanations when the user is nervous. Furthermore, the analysis unit can use generative AI to quickly analyze the video and provide concise explanations when the user is relaxed. Further, the analysis unit can use generative AI to focus on key points when the user is anxious. Thus, by adjusting the video analysis method according to the user's emotions, more appropriate analysis results can be provided. Emotion inference can be achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generating AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the analysis unit can also be implemented using AI, or AI can be omitted. For example, the analysis unit can input the user's facial expression data into the generative AI, which will then perform emotion inference.
[0097] The analysis unit can detect specific actions or gestures while analyzing videos and refine the analysis results accordingly. For example, the analysis unit can use generative AI to detect hand movements in videos and analyze operation steps in detail. Generative AI can detect specific actions or gestures based on video content. For example, the analysis unit can use generative AI to detect facial expressions in videos and analyze user intentions. Furthermore, the analysis unit can use generative AI to detect body movements in videos and analyze work processes. Thus, by detecting specific actions or gestures, the analysis results can be refined. Generative AI can, for example, use computer vision technology to detect specific actions or gestures in videos with high precision. Generative AI can learn from large amounts of video data to achieve high-precision action detection. For example, generative AI can detect hand movements or facial expressions in videos with high precision and analyze their content. Therefore, the analysis unit can use generative AI to detect specific actions or gestures while analyzing videos and refine the analysis results accordingly.
[0098] The analysis unit can analyze background and ambient sounds while analyzing video, and add information about the work environment. For example, the analysis unit can use generative AI to analyze background sounds in the video and assess the noise level of the work environment. Generative AI can analyze background and ambient sounds based on video content. For example, the analysis unit can use generative AI to analyze ambient sounds in the video to determine the work location. Furthermore, the analysis unit can use generative AI to analyze speech in the video and extract work instructions. Thus, by analyzing background and ambient sounds, information about the work environment can be added. For example, generative AI can use speech recognition technology to analyze background and ambient sounds in the video with high accuracy. Generative AI can learn from large amounts of speech data to achieve high-precision speech analysis. For example, generative AI can analyze background and ambient sounds in the video with high accuracy and evaluate their content. Therefore, the analysis unit can use generative AI to analyze background and ambient sounds while analyzing video and add information about the work environment.
[0099] The analysis unit can infer the user's emotions and determine the priority of analysis results based on the inferred emotions. For example, the analysis unit can utilize generative AI to infer the user's emotions. Generative AI can analyze the user's facial expressions and speech to infer emotions. For instance, the analysis unit can use generative AI to prioritize displaying important analysis results when the user is stressed. Furthermore, the analysis unit can use generative AI to display detailed analysis results sequentially when the user is relaxed. Further, the analysis unit can use generative AI to quickly display analysis results highlighting key points when the user is anxious. Thus, by prioritizing analysis results based on the user's emotions, important information can be provided first. Emotion inference can be achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text generation AI (such as LLM) or multimodal generation AI, but is not limited to these. Some or all of the above processing in the analysis unit can also be implemented using AI, or AI can be omitted. For example, the analysis unit can input the user's facial expression data into the generative AI, which will then perform emotion inference.
[0100] The parsing department can improve parsing accuracy by referencing the user's past video parsing history. For example, the parsing department can utilize generative AI to reference the user's past video parsing history. Generative AI can optimize the parsing algorithm based on past parsing history. For instance, the parsing department can use generative AI to refer to the user's past parsing history to improve the accuracy when parsing similar videos. Furthermore, the parsing department can use generative AI to optimize the parsing algorithm based on the user's past parsing results. Further, the parsing department can use generative AI to analyze the user's past parsing history to grasp parsing trends. Thus, by referring to past parsing history, parsing accuracy can be improved. For example, generative AI can use machine learning algorithms to learn from the user's past parsing history. Generative AI can optimize the parsing algorithm based on a large amount of parsing history data. For example, generative AI can analyze the user's past parsing history with high precision, improving parsing accuracy. Therefore, the parsing department can improve parsing accuracy by referencing the user's past video parsing history using generative AI.
[0101] The analysis unit can consider the user's geographic location information to customize the analysis results when analyzing video. For example, the analysis unit can utilize generative AI to consider the user's geographic location information. Generative AI can customize analysis results based on the user's location information. For example, the analysis unit can use generative AI to analyze regionally specific work procedures based on the user's geographic location information. Furthermore, the analysis unit can use generative AI to refer to the user's location information to provide analysis results that conform to the regional language and culture. Further, the analysis unit can use generative AI to perform analysis considering regional environmental conditions based on the user's location information. Thus, by considering geographic location information, the analysis results can be customized. For example, generative AI can utilize location information services to obtain the user's geographic location information. Generative AI can optimize the analysis results based on location information data. For example, generative AI can obtain the user's location information with high accuracy and customize the analysis results based on this information. Therefore, the analysis unit can use generative AI to consider the user's geographic location information to customize the analysis results when analyzing video.
[0102] The generation unit can infer the user's emotions and adjust the presentation of generated images and descriptions based on the inferred emotions. For example, the generation unit can utilize generative AI to infer user emotions. Generative AI can analyze user facial expressions and speech to infer emotions. For instance, the generation unit can use generative AI to generate concise and highly visual images and descriptions when the user is nervous. Furthermore, the generation unit can use generative AI to generate detailed and colorful images and descriptions when the user is relaxed. Further, the generation unit can use generative AI to generate concise descriptions and images highlighting key points when the user is anxious. Thus, by adjusting the presentation of images and descriptions according to the user's emotions, a more appropriate user manual can be provided. Emotion inference can be achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text generation AI (such as LLM) or multimodal generation AI, but is not limited to these. Some or all of the above processing in the generation unit can also be achieved through AI, or AI can be omitted. For example, the generation unit can input the user's facial expression data into the generative AI, which will then perform emotion inference.
[0103] The generation unit can generate images and explanatory text to emphasize specific scenes or important steps in a video during the generation process. For example, the generation unit can utilize generative AI to generate images and explanatory text emphasizing important operational steps in the video. Generative AI can emphasize specific scenes or important steps based on the video content. For example, the generation unit can use generative AI to capture specific scenes in the video and attach detailed explanatory text. Furthermore, the generation unit can use generative AI to highlight important points in the video for visual emphasis. Thus, by emphasizing specific scenes or important steps, an easy-to-understand operation manual can be provided. For example, generative AI can utilize computer vision technology to detect specific scenes or important steps in the video with high precision. Generative AI can learn from large amounts of video data to achieve high-precision scene detection. For example, generative AI can detect important operational steps in a video with high precision and emphasize their content. Therefore, the generation unit can use generative AI to generate images and explanatory text to emphasize specific scenes or important steps in the video during the generation process.
[0104] The generation department can generate descriptive text containing interactive elements based on video content during the generation process. For example, the generation department can utilize generative AI to generate descriptive text containing interactive elements (such as clickable links or buttons) based on video content. Generative AI can generate descriptive text containing interactive elements based on video content. For example, the generation department can use generative AI to add links to specific operation steps in the video, providing detailed explanations. Furthermore, the generation department can use generative AI to set buttons in important scenes in the video, displaying relevant information. Further, the generation department can use generative AI to add interactive elements to each step in the video, allowing users to access detailed information. Thus, by generating descriptive text containing interactive elements, users can more easily access detailed information. Generative AI can, for example, utilize web technologies to generate descriptive text containing interactive elements. Generative AI can learn from large amounts of web data to generate highly accurate interactive elements. For example, generative AI can add links or buttons to specific operation steps in the video with high accuracy, realizing content interactivity. Thus, the generation department can generate descriptive text containing interactive elements based on video content using generative AI.
[0105] The generation unit can infer the user's emotions and adjust the length of the generated images and descriptions based on the inferred emotions. For example, the generation unit can utilize generative AI to infer the user's emotions. Generative AI can analyze the user's facial expressions and speech to infer emotions. For instance, the generation unit can use generative AI to generate concise and key-point descriptions and images when the user is anxious. Furthermore, the generation unit can use generative AI to generate detailed descriptions and images when the user is relaxed. Further, the generation unit can use generative AI to generate descriptions and images with visual stimulating effects when the user is excited. Thus, by adjusting the length of images and descriptions according to the user's emotions, a more appropriate user manual can be provided. Emotion inference can be achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text generation AI (such as LLM) or multimodal generation AI, but is not limited to these. Some or all of the above processing in the generation unit can also be implemented using AI, or AI can be omitted. For example, the generation unit can input the user's facial expression data into the generative AI, which will then perform emotion inference.
[0106] The content generation department can customize generated content by referencing the user's past usage history of the user manual during the generation process. For example, the generation department can utilize generative AI to reference the user's past usage history of the user manual. Generative AI can customize generated content based on past usage history. For example, the generation department can use generative AI to reference the user's past usage history of the user manual and customize similar content. Furthermore, the generation department can use generative AI to generate optimal explanatory text and images based on the user's past usage history. Further, the generation department can use generative AI to analyze the user's past usage history and generate user manuals that match user preferences. Thus, by referencing past usage history, the generated content can be customized. For example, generative AI can use machine learning algorithms to learn the user's past usage history. Generative AI can optimize generated content based on a large amount of historical usage data. For example, generative AI can analyze the user's past usage history with high precision and customize generated content. Therefore, the generation department can use generative AI to customize generated content by referencing the user's past usage history of the user manual during the generation process.
[0107] The generation unit can consider the user's device information during generation to generate optimal images and descriptions. For example, the generation unit can utilize generative AI to consider the user's device information. Generative AI can generate optimal images and descriptions based on the user's device information. For example, the generation unit can use generative AI to generate images and descriptions suitable for the user's device screen size. Furthermore, the generation unit can use generative AI to consider the performance of the user's device and generate images and descriptions in the optimal format. Further, the generation unit can use generative AI to refer to the user's device usage and select the optimal display method. Thus, by considering device information, optimal images and descriptions can be generated. For example, generative AI can utilize device information services to obtain the user's device information. Generative AI can optimize the generated content based on device information data. For example, generative AI can obtain the user's device information with high accuracy and generate optimal images and descriptions based on that information. Thus, the generation unit can use generative AI to consider the user's device information during generation to generate optimal images and descriptions.
[0108] The production department can infer the user's emotions and adjust the structure of the user manual based on these inferred emotions. For example, the production department can utilize generative AI to infer user emotions. Generative AI can analyze user facial expressions and speech to infer emotions. For instance, the production department can use generative AI to create a concise and highly visual user manual when the user is nervous. Furthermore, the production department can use generative AI to create a detailed and colorful user manual when the user is relaxed. Further, the production department can use generative AI to create a concise user manual highlighting key points when the user is anxious. Thus, by adjusting the structure of the user manual according to the user's emotions, a more appropriate user manual can be provided. Emotion inference can be achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generating AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the production department can also be achieved through AI, or AI can be omitted. For example, the production department can input the user's facial expression data into the generative AI, which will then perform emotion inference.
[0109] The production department can reference past user feedback when creating user manuals to identify areas for improvement. For example, the production department can utilize generative AI to reference past user feedback. Generative AI can then use this feedback to identify areas for improvement in the user manual. For instance, the production department can use generative AI to reference past user feedback and create user manuals that reflect these improvements. Furthermore, the production department can use generative AI to optimize the content of the user manual based on user feedback. Moreover, the production department can use generative AI to analyze past user feedback and create user manuals that meet user needs. Thus, by referencing past feedback, the production department can identify areas for improvement in the user manual. For example, generative AI can use machine learning algorithms to learn from past user feedback. Generative AI can optimize user manuals based on a large amount of feedback data. For instance, generative AI can analyze past user feedback with high precision to identify areas for improvement in the user manual. Therefore, the production department can use generative AI to reference past user feedback when creating user manuals to identify areas for improvement.
[0110] The production department can use templates customized based on video content when creating operation manuals. For example, the production department can utilize generative AI to use templates customized based on video content. Generative AI can select the optimal template based on video content. For example, the production department can use generative AI to analyze video content and select the optimal template to create the operation manual. Furthermore, the production department can use generative AI to create operation manuals using corresponding templates for each step of the video. Further, the production department can use generative AI to generate customized templates based on video content to create operation manuals. Thus, by using customized templates, more suitable operation manuals can be provided. For example, generative AI can use template generation algorithms to generate customized templates based on video content. Generative AI can learn from a large amount of template data to generate high-precision templates. For example, generative AI can analyze video content with high precision and generate the optimal template based on that content. Therefore, the production department can use templates customized based on video content when creating operation manuals using generative AI.
[0111] The production department can infer users' emotions and prioritize user manuals based on these inferred emotions. For example, the production department can utilize generative AI to infer user emotions. Generative AI can analyze users' facial expressions and speech to infer emotions. For instance, the production department can use generative AI to prioritize creating important user manuals when users are stressed. Furthermore, the production department can use generative AI to create detailed user manuals sequentially when users are relaxed. Further, the production department can use generative AI to quickly create user manuals highlighting key points when users are anxious. Thus, by prioritizing user manuals based on their emotions, important user manuals can be provided first. Emotion inference can be achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generating AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the production department can also be achieved through AI, or AI can be omitted. For example, the production department can input users' facial expression data into generative AI, which will then perform emotion inference.
[0112] The production department can consider the user's geographical location information when creating user manuals to produce the optimal manual. For example, the production department can utilize generative AI to consider the user's geographical location information. Generative AI can create the optimal user manual based on the user's location information. For instance, the production department can use generative AI to create a user manual that includes region-specific operating procedures based on the user's geographical location information. Furthermore, the production department can use generative AI to create a user manual that conforms to the region's language and culture, referencing the user's location information. Further, the production department can use generative AI to create a user manual that considers the region's environmental conditions based on the user's location information. Thus, by considering geographical location information, the optimal user manual can be created. For example, generative AI can utilize location information services to obtain the user's geographical location information. Generative AI can optimize the content of the user manual based on location information data. For example, generative AI can obtain the user's location information with high accuracy and create the optimal user manual based on this information. Therefore, the production department can use generative AI to consider the user's geographical location information when creating user manuals to produce the optimal manual.
[0113] The production department can analyze users' social media activities when creating user manuals and incorporate relevant information into the manuals. For example, the production department can utilize generative AI to analyze users' social media activities. Generative AI can incorporate relevant information based on social media activities into the user manual. For example, the production department can use generative AI to analyze users' social media activities and reflect this information in the user manual. Furthermore, the production department can use generative AI to optimize the user manual content by referencing user feedback on social media. Further, the production department can use generative AI to create user manuals that match user preferences based on users' social media activities. Thus, by analyzing social media activities, relevant information can be incorporated into the user manual. For example, generative AI can utilize social media analysis algorithms to analyze users' social media activities. Generative AI can learn from large amounts of social media data to achieve high-precision analysis. For example, generative AI can analyze users' social media activities with high precision and reflect this information in the user manual. Therefore, the production department can use generative AI to analyze users' social media activities when creating user manuals and incorporate relevant information into the manuals.
[0114] The support department can infer the user's emotions and adjust the chatbot's responses based on these inferred emotions. For example, the support department can utilize generative AI to infer user emotions. Generative AI can analyze a user's facial expressions and speech to infer emotions. For instance, the support department can use generative AI to provide a calm response from the chatbot when the user is nervous. Furthermore, the support department can use generative AI to provide a friendly response from the chatbot when the user is relaxed. Further, the support department can use generative AI to provide a quick and concise response from the chatbot when the user is anxious. Thus, by adjusting the chatbot's responses based on the user's emotions, more appropriate responses can be provided. Emotion inference can be achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generating AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the support department can also be implemented using AI, or AI can be omitted. For example, the support department can input the user's facial expression data into the generative AI, which will then perform emotion inference.
[0115] The support department can refer to a user's past question history to provide the optimal response when the chatbot answers. For example, the support department can utilize generative AI to reference the user's past question history. Generative AI can provide the optimal response based on past question history. For example, the support department can use generative AI to refer to a user's past question history to provide the optimal response to similar questions. Furthermore, the support department can use generative AI based on the user's past question history to improve the accuracy of the response. Further, the support department can use generative AI to analyze the user's past question history to provide responses that match the user's preferences. Thus, by referring to past question history, the optimal response can be provided. For example, generative AI can use question history analysis algorithms to analyze the user's past question history. Generative AI can learn from a large amount of question history data to achieve high-precision analysis. For example, generative AI can analyze a user's past question history with high precision and provide the optimal response based on this information. Therefore, the support department can use generative AI to refer to the user's past question history to provide the optimal response when the chatbot answers.
[0116] The support department can provide customized responses based on the user's current situation or environment when the chatbot replies. For example, the support department can utilize generative AI to analyze the user's current situation or environment. Generative AI can provide the optimal response based on the user's situation or environment. For example, the support department can use generative AI to analyze the user's current situation and provide the optimal response. Furthermore, the support department can use generative AI to refer to the user's environmental information and provide responses suitable for the environment. Further, the support department can use generative AI to provide customized responses based on the user's current situation. Thus, by providing customized responses based on the current situation or environment, more appropriate responses can be provided. For example, generative AI can use situation analysis algorithms to analyze the user's current situation. Generative AI can learn from large amounts of situation data to achieve high-precision analysis. For example, generative AI can accurately analyze the user's current situation and provide the optimal response based on this information. Therefore, the support department can use generative AI to provide customized responses based on the user's current situation or environment when the chatbot replies.
[0117] The support department can infer the user's emotions and determine the priority of chatbot responses based on the inferred emotions. For example, the support department can utilize generative AI to infer user emotions. Generative AI can analyze user facial expressions and speech to infer emotions. For instance, the support department can use generative AI to prioritize important responses when the user is stressed. Furthermore, the support department can use generative AI to provide detailed responses sequentially when the user is relaxed. Further, the support department can use generative AI to quickly provide concise responses when the user is anxious. Thus, by prioritizing responses based on user emotions, important responses can be provided first. Emotion inference can be achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generating AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the support department can also be implemented using AI, or AI can be omitted. For example, the support department can input the user's facial expression data into the generative AI, which will then perform emotion inference.
[0118] The support department can consider the user's geographic location information to provide the optimal response when the chatbot replies. For example, the support department can utilize generative AI to consider the user's geographic location information. Generative AI can provide the optimal response based on the user's location information. For example, the support department can use generative AI to provide responses containing region-specific information based on the user's geographic location information. Furthermore, the support department can use generative AI to provide responses that conform to the region's language and culture, taking into account the user's location information. Further, the support department can use generative AI to provide responses that consider the region's environmental conditions based on the user's location information. Thus, by considering geographic location information, the optimal response can be provided. For example, generative AI can utilize location information services to obtain the user's geographic location information. Generative AI can optimize the response content based on location information data. For example, generative AI can obtain the user's location information with high accuracy and provide the optimal response based on that information. Therefore, the support department can use generative AI to consider the user's geographic location information to provide the optimal response when the chatbot replies.
[0119] The support department can analyze users' social media activity and provide relevant information when the chatbot responds. For example, the support department can utilize generative AI to analyze users' social media activity. Generative AI can provide relevant information based on social media activity. For instance, the support department can use generative AI to analyze users' social media activity and reflect this information in the response. Furthermore, the support department can use generative AI to optimize the response content by referencing user feedback on social media. Further, the support department can use generative AI to provide responses that match user preferences based on users' social media activity. Thus, by analyzing social media activity, relevant information can be provided. For example, generative AI can utilize social media analysis algorithms to analyze users' social media activity. Generative AI can learn from large amounts of social media data to achieve high-precision analysis. For example, generative AI can analyze users' social media activity with high precision and provide the optimal response based on this information. Therefore, the support department can use generative AI to analyze users' social media activity and provide relevant information when the chatbot responds.
[0120] The system involved in this embodiment is not limited to the above examples. For example, various modifications can be made as follows.
[0121] The analysis unit can infer the user's emotions and adjust the display of the analysis results accordingly. For example, when the user is stressed, the analysis unit will emphasize important information and display it concisely. When the user is relaxed, detailed information can be displayed sequentially. Furthermore, when the user is anxious, key points can be quickly displayed. Thus, by adjusting the display of analysis results based on the user's emotions, information can be provided more appropriately.
[0122] The generation department can infer the user's emotions and adjust the style of the generated images and descriptions accordingly. For example, when the user is tense, concise and highly visual images and descriptions are generated. When the user is relaxed, detailed and colorful images and descriptions are generated. Furthermore, when the user is anxious, concise descriptions and images highlighting key points are generated. Thus, by adjusting the style of images and descriptions based on the user's emotions, more appropriate user manuals can be provided.
[0123] The production department can infer the user's emotions and adjust the structure of the user manual accordingly. For example, when the user is nervous, a concise and highly visual user manual can be created. When the user is relaxed, a detailed and colorful user manual can be created. Furthermore, when the user is anxious, a short user manual highlighting key points can be created. Thus, by adjusting the structure of the user manual based on their emotions, a more appropriate user manual can be provided.
[0124] The support department can infer a user's emotions and adjust the chatbot's responses accordingly. For example, when a user is nervous, the chatbot responds in a calm tone. When the user is relaxed, it responds in a friendly tone. Furthermore, when the user is anxious, it provides a quick and concise response. Thus, by adjusting the chatbot's responses based on the user's emotions, more appropriate answers can be provided.
[0125] The analysis unit can infer the user's emotions and determine the priority of analysis results based on the inferred emotions. For example, when the user is stressed, important analysis results are displayed first. When the user is relaxed, detailed analysis results can be displayed sequentially. Furthermore, when the user is anxious, analysis results highlighting key points can be quickly displayed. Thus, by prioritizing analysis results based on the user's emotions, important information can be provided first.
[0126] The analysis unit can detect specific actions or gestures while analyzing videos and refine the analysis results accordingly. For example, generative AI can be used to detect hand movements in videos and analyze operation steps in detail. Furthermore, facial expressions in videos can be detected to analyze user intentions. Even further, body movements in videos can be detected to analyze work processes. Thus, by detecting specific actions or gestures, the analysis results can be refined.
[0127] The analysis unit can analyze background and ambient sounds while analyzing video, and add information about the work environment. For example, generative AI can be used to analyze background sounds in the video to assess the noise level of the work environment. Furthermore, ambient sounds in the video can be analyzed to determine the work location. Further, speech in the video can be analyzed to extract work instructions. Thus, by analyzing background and ambient sounds, information about the work environment can be added.
[0128] The parsing department can improve parsing accuracy by referencing a user's past video parsing history. For example, generative AI can be used to improve the accuracy of parsing similar videos by referencing the user's past parsing history. Furthermore, the parsing algorithm can be optimized based on the user's past parsing results. Moreover, the parsing trend can be identified by analyzing the user's past parsing history. Therefore, by referring to past parsing history, parsing accuracy can be improved.
[0129] The generation unit can generate images and explanatory text to emphasize specific scenes or important steps in a video during the generation process. For example, generative AI can be used to generate images and explanatory text that highlight important operational steps in a video. Furthermore, specific scenes in the video can be captured and detailed explanations can be added. Additionally, important points in the video can be highlighted for visual emphasis. Thus, by emphasizing specific scenes or important steps, an easy-to-understand instruction manual can be provided.
[0130] The generation department can generate descriptive text containing interactive elements based on video content during the generation process. For example, generative AI can be used to generate descriptive text with interactive elements (such as clickable links or buttons) based on video content. Furthermore, links can be added to specific operation steps in the video, providing detailed explanations. Additionally, buttons can be placed in important scenes within the video to display relevant information. Thus, by generating descriptive text with interactive elements, users can more easily access detailed information.
[0131] The following is a brief description of the processing flow of Implementation Method 2.
[0132] Step 1: The analysis unit analyzes the video. The analysis unit uses generative AI to analyze the video content, examining each frame and detecting specific objects or actions. For example, it detects specific operational steps in the video and generates corresponding images and descriptions.
[0133] Step 2: The generation unit generates images and explanatory text based on the video content parsed by the parsing unit. The generation unit uses generative AI to generate images and explanatory text for each step, generating images and explanatory text corresponding to specific operation steps in the video.
[0134] Step 3: The production department creates an operation manual based on the images and explanatory text generated by the generation department. The production department uses generative AI to automatically determine the structure and format of the operation manual based on the generated images and explanatory text, arranging the images and explanatory text for each step in sequence to create a clear and concise operation manual.
[0135] Step 4: The support department has set up a chatbot to answer questions in the operation manuals produced by the production department. Utilizing AI, the chatbot responds to user questions about the manual's content, providing detailed explanations for specific operational steps.
[0136] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires voice representing the user's input to the result of the specific processing. The control unit 46A sends the voice data representing the user's input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0137] Data generation model 58 is what is known as generative AI (Artificial Intelligence). An example of data generation model 58 includes ChatGPT (registered trademark) (Internet search).<URL:https: / / openai.com / blog / chatgpt> Generative AI, such as data generation model 58, is obtained by deep learning through a neural network. The data generation model 58 is input with a prompt containing instructions, and with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data of still images or data of moving images). The data generation model 58 infers the input inference data according to the instructions shown in the prompt, and outputs the inference result in one or more data forms such as speech data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the above-described specific processing while using the data generation model 58. The data generation model 58 can also be a fine-tuned model to output inference results from prompts without instructions; in this case, the data generation model 58 can output inference results from prompts without instructions. The data processing device 12, etc., includes multiple data generation models 58, including AI other than generative AI. AI other than generative AI includes, but is not limited to, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes. Furthermore, AI can also act as an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to this example. Moreover, processing performed by AI, including generative AI, can be replaced by rule-based processing, and vice versa.
[0138] Furthermore, the processing performed by the aforementioned data processing system 10 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0139] Each of the elements comprising the aforementioned analysis unit, generation unit, production unit, and support unit can, for example, be implemented in at least one of the smart device 14 and the data processing device 12. For example, the analysis unit may use the camera 42 of the smart device 14 to capture video, and the specific processing unit 290 of the data processing device 12 may analyze the video content. The generation unit may, for example, generate images and explanatory text based on the content analyzed by the specific processing unit 290 of the data processing device 12. The production unit may, for example, create an operation manual based on the images and explanatory text generated by the specific processing unit 290 of the data processing device 12. The support unit may, for example, utilize a chatbot set up by the control unit 46A of the smart device 14 to answer user questions. The correspondence between each unit and the device or control unit is not limited to the above examples and can be modified in various ways.
[0140] Second Implementation Method
[0141] Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0142] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. One example of the data processing device 12 is a server.
[0143] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.
[0144] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0145] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.
[0146] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, used to capture the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).
[0147] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.
[0148] Figure 4 An example of the main functions of the data processing device 12 and the smart glasses 214 is shown. Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.
[0149] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.
[0150] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes inferring and predicting the user's emotions, performing various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).
[0151] In the smart glasses 214, specific processing is performed by the processor 46. A specific processing program 60 is stored in the memory 50. The processor 46 reads the specific processing program 60 from the memory 50 and executes the read specific processing program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific processing program 60 executed on the RAM 48. Furthermore, the smart glasses 214 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.
[0152] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).
[0153] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires voice input representing the user's input to the specific processing result. The control unit 46A sends the voice data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0154] Data generation model 58 is a so-called generative AI. An example of data generation model 58 includes generative AIs such as ChatGPT. Data generation model 58 is obtained by performing deep learning on a neural network. A prompt containing instructions is input to data generation model 58, along with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data from still images or moving images). Data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. Specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. Data generation model 58 can also be a fine-tuned model to output inference results from prompts without instructions; in this case, data generation model 58 can output inference results from prompts without instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.
[0155] The data processing system 210 of the second embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0156] Each of the elements comprising the aforementioned analysis unit, generation unit, production unit, and support unit can, for example, be implemented in at least one of the smart glasses 214 and the data processing device 12. For example, the analysis unit may use the camera 42 of the smart glasses 214 to capture video, and the specific processing unit 290 of the data processing device 12 may analyze the video content. The generation unit may, for example, generate images and explanatory text based on the content analyzed by the specific processing unit 290 of the data processing device 12. The production unit may, for example, create an operation manual based on the images and explanatory text generated by the specific processing unit 290 of the data processing device 12. The support unit may, for example, use a chatbot set up by the control unit 46A of the smart glasses 214 to answer user questions. The correspondence between each unit and the device or control unit is not limited to the above examples and can be modified in various ways.
[0157] Third Implementation Method
[0158] Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0159] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. One example of the data processing device 12 is a server.
[0160] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.
[0161] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0162] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.
[0163] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, used to capture the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).
[0164] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.
[0165] Figure 6 An example of the main functions of the data processing device 12 and the head-mounted terminal 314 is shown. Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.
[0166] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.
[0167] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes inferring and predicting the user's emotions, performing various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).
[0168] In the head-mounted terminal 314, specific processing is performed by the processor 46. A specific program 60 is stored in the memory 50. The processor 46 reads the specific program 60 from the memory 50 and executes the read specific program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific program 60 executed on the RAM 48. Furthermore, the head-mounted terminal 314 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.
[0169] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).
[0170] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires voice representing the user's input to the specific processing result. The control unit 46A sends the voice data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0171] Data generation model 58 is a so-called generative AI. An example of data generation model 58 includes generative AIs such as ChatGPT. Data generation model 58 is obtained by performing deep learning on a neural network. A prompt containing instructions is input to data generation model 58, along with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data from still images or moving images). Data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. Specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. Data generation model 58 can also be a fine-tuned model to output inference results from prompts without instructions; in this case, data generation model 58 can output inference results from prompts without instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.
[0172] The data processing system 310 of the third embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0173] Each of the elements comprising the aforementioned analysis unit, generation unit, production unit, and support unit can, for example, be implemented in at least one of the head-mounted terminal 314 and the data processing device 12. For example, the analysis unit may use the camera 42 of the head-mounted terminal 314 to capture video, and the specific processing unit 290 of the data processing device 12 may analyze the video content. The generation unit may, for example, generate images and explanatory text based on the content analyzed by the specific processing unit 290 of the data processing device 12. The production unit may, for example, create an operation manual based on the images and explanatory text generated by the specific processing unit 290 of the data processing device 12. The support unit may, for example, use a chatbot set up by the control unit 46A of the head-mounted terminal 314 to answer user questions. The correspondence between each unit and the device or control unit is not limited to the above examples and can be modified in various ways.
[0174] Fourth Implementation Method
[0175] Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0176] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0177] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.
[0178] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control object 443. Computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and control object 443 are also connected to the bus 52.
[0179] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.
[0180] The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, used to photograph the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).
[0181] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.
[0182] The controlled object 443 includes a display device, LEDs for the eyes, and motors for driving the arms, hands, and feet. The posture and movements of the robot 414 are controlled by controlling the motors for the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, facial expressions of the robot 414 can also be expressed by controlling the illumination state of the LEDs for the robot 414's eyes.
[0183] Figure 8 An example of the main functions of the data processing device 12 and the robot 414 is shown. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.
[0184] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.
[0185] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes inferring and predicting the user's emotions, performing various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).
[0186] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in memory 50. Processor 46 reads the specific program 60 from memory 50 and executes the read specific program 60 on RAM 48. Specific processing is achieved by processor 46 acting as control unit 46A based on the specific program 60 executed on RAM 48. Furthermore, robot 414 may also have the same data generation model and emotion-specific model as data generation model 58 and emotion-specific model 59, and use these models to perform the same processing as specific processing unit 290.
[0187] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).
[0188] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires voice representing the user's input regarding the result of the specific processing. The control unit 46A sends the voice data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0189] Data generation model 58 is a so-called generative AI. An example of data generation model 58 includes generative AIs such as ChatGPT. Data generation model 58 is obtained by performing deep learning on a neural network. A prompt containing instructions is input to data generation model 58, along with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data from still images or moving images). Data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. Specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. Data generation model 58 can also be a fine-tuned model to output inference results from prompts without instructions; in this case, data generation model 58 can output inference results from prompts without instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.
[0190] The data processing system 410 of the fourth embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0191] Each of the elements comprising the aforementioned analysis unit, generation unit, production unit, and support unit can, for example, be implemented in at least one of the robot 414 and the data processing device 12. For example, the analysis unit may use the robot 414's camera 42 to capture video, and the video content may be analyzed by a specific processing unit 290 of the data processing device 12. The generation unit may, for example, generate images and explanatory text based on the content analyzed by the specific processing unit 290 of the data processing device 12. The production unit may, for example, create an operation manual based on the images and explanatory text generated by the specific processing unit 290 of the data processing device 12. The support unit may, for example, use a chatbot set up by the control unit 46A of the robot 414 to answer user questions. The correspondence between the various units and the device or control unit is not limited to the above examples and can be modified in various ways.
[0192] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine the user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine the user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The robot's emotions can be determined by the emotion-specific model 59. In addition, the emotion-specific model 59 can also determine the robot's emotions in the same way, and the specific processing unit 290 can also perform specific processing using the robot's emotions.
[0193] Figure 9 This is a diagram representing an emotion map 400 that maps various emotions. In the emotion map 400, emotions are arranged radially from the center in concentric circles. The closer to the center of the concentric circles, the more primitive the emotion is. Further out on the concentric circles, emotions are arranged representing states or actions arising from mood. Emotion is a concept that includes both feelings and mental states. To the left of the concentric circles, emotions generated by reactions occurring in the brain are arranged roughly. To the right of the concentric circles, emotions guided by situational judgments are arranged roughly. Above and below the concentric circles, emotions generated by reactions occurring in the brain and guided by situational judgments are arranged roughly. Furthermore, the emotion of "pleasure" is arranged above the concentric circles, and the emotion of "unpleasantness" is arranged below. Thus, in the emotion map 400, various emotions are mapped according to the structure of emotion generation, while easily generated emotions are mapped nearby.
[0194] These emotions are distributed at the 3 o'clock position on the Emotion Chart 400, and usually fluctuate between peace and unease. In the right half of the Emotion Chart 400, because situational awareness is more dominant than internal feelings, it gives a sense of calm.
[0195] The inner side of the emotion diagram 400 represents the mind, and the outer side of the emotion diagram 400 represents actions. Therefore, the further you go to the outer side of the emotion diagram 400, the more the emotion can be seen (manifested in actions).
[0196] Here, human emotions are based on a balance of various factors such as posture and blood sugar levels. When these balances deviate from the ideal, it indicates unhappiness; when they approach the ideal, it indicates pleasure. In robots, cars, and motorcycles, emotions can also be created based on a balance of factors such as posture and remaining battery power. When these balances deviate from the ideal, it indicates unhappiness; when they approach the ideal, it indicates pleasure. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Speech Emotion Recognition and Brain Physiological Signal Analysis Systems for Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the "Reaction" domain, where sensation is dominant, are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the "Situation" domain, where situational cognition is dominant, are arranged.
[0197] The emotion map defines two types of emotions that promote learning. One is a negative emotion located near the middle of "repentance" or "reflection" on the situation side. That is, when the robot experiences negative emotions such as "I never want to feel this way again" or "I never want to be scolded again." The other is a positive emotion located near "desire" on the response side. That is, when the robot experiences positive feelings such as "wanting more" or "wanting to know more."
[0198] The emotion-specific model 59 feeds user input into a pre-learned neural network to obtain emotion values representing each emotion shown in the emotion graph 400, and determines the user's emotion. This neural network is pre-learned based on multiple learning data sets that combine user input with emotion values representing each emotion shown in the emotion graph 400. Furthermore, this neural network is learned to... Figure 10 As shown in sentiment graph 900, sentiment values in nearby configurations are similar to each other. Figure 10 Examples show that multiple emotions such as "peace of mind", "stability", and "reassurance" have similar emotional values.
[0199] In the above embodiments, a specific processing is described by a single computer 22, but the technology disclosed herein is not limited to this, and distributed processing by multiple computers, including computer 22, is also possible.
[0200] In the above embodiments, an example of storing a specific processing program 56 in memory 32 is illustrated, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 performs specific processing according to the specific processing program 56.
[0201] Alternatively, the specific processing program 56 can be stored in a storage device such as a server connected to the data processing device 12 via a network 54, and the specific processing program 56 can be downloaded and installed into the computer 22 upon request from the data processing device 12.
[0202] Furthermore, it is not necessary to store the entire specific process 56 in a storage device such as a server connected to the data processing device 12 via the network 54, nor is it necessary to store the entire specific process 56 in the memory 32; a portion of the specific process 56 may also be stored.
[0203] As a hardware resource for performing specific processing, various processors can be used. For example, a CPU is a general-purpose processor that functions as a hardware resource for performing specific processing by executing software, i.e., programs. Additionally, a dedicated circuit can be listed as a processor; it is a processor with a circuit structure specifically designed for performing specific processing, such as a FPGA (Field-Programmable Gate Array), PLD (Programmable Logic Device), or ASIC (Application Specific Integrated Circuit). Every processor has built-in or connected memory, and every processor executes specific processing by using that memory.
[0204] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Furthermore, the hardware resources for performing a specific process can also be a single processor.
[0205] As an example of a single processor, the first type consists of a combination of one or more CPUs and software, which functions as a hardware resource to perform specific processing. The second type uses a processor, such as a System-on-a-chip (SoC), which implements the entire system functionality, including multiple hardware resources for performing specific processing, using a single IC chip. In this case, the specific processing is implemented using one or more of the aforementioned processors that serve as hardware resources.
[0206] Furthermore, as the hardware architecture of these various processors, more specifically, circuits combining semiconductor elements and other circuit components can be used. Moreover, the specific process described above is merely an example. Therefore, it goes without saying that, without departing from the main point, unnecessary steps can be removed, new steps can be added, or the processing order can be changed.
[0207] Furthermore, although the above examples have been described in terms of first to fourth embodiments, some or all of these embodiments can be combined. Additionally, the smart device 14, smart glasses 214, head-mounted terminal 314, and robot 414 are just examples and can be combined separately, or other devices may be used. Furthermore, although the above examples have been described in terms of morphological example 1 and morphological example 2, these can also be combined.
[0208] The foregoing descriptions and illustrations are detailed explanations of the parts covered by this disclosure and are merely one example of this disclosure. For instance, the descriptions of the above-described structure, function, role, and effect are just one example of the structure, function, role, and effect of the parts covered by this disclosure. Therefore, it goes without saying that, without departing from the spirit of this disclosure, unnecessary parts can be deleted, new elements can be added, or replacements can be made to the foregoing descriptions and illustrations. Furthermore, to avoid confusion and facilitate understanding of the parts covered by this disclosure, explanations of technical common sense that does not require special explanation for implementing this disclosure have been omitted from the foregoing descriptions and illustrations.
[0209] All documents, patent applications and technical standards described in this specification are incorporated herein by reference as if they were specifically and individually described as incorporated by reference.
[0210] [Postscript 1]
[0211] A system, characterized in that it comprises:
[0212] The parsing unit is used to parse video.
[0213] The generation unit is used to generate images and explanatory text based on the video content parsed by the parsing unit;
[0214] The production department is used to create operation manuals based on the images and explanatory texts generated by the generation department.
[0215] The support department is equipped with a chatbot for responding to operation manuals produced by the production department.
[0216] [Postscript 2]
[0217] The system as described in Appendix 1 is characterized in that,
[0218] The analysis unit analyzes video content using generative AI.
[0219] [Postscript 3]
[0220] The system as described in Appendix 1 is characterized in that,
[0221] The generation unit generates images and explanatory text for each step using generative AI.
[0222] [Postscript 4]
[0223] The system as described in Appendix 1 is characterized in that,
[0224] The production department creates operation manuals based on images and explanatory texts generated by generative AI.
[0225] [Postscript 5]
[0226] The system as described in Appendix 1 is characterized in that,
[0227] When users ask questions about the contents of the user manual, the support department uses a chatbot to answer them.
[0228] [Postscript 6]
[0229] The system as described in Appendix 1 is characterized in that,
[0230] The support department uses generative AI for version management, ensuring that it always provides the latest user manuals.
[0231] [Postscript 7]
[0232] The system as described in Appendix 1 is characterized in that,
[0233] The analysis unit infers the user's emotions and adjusts the video analysis method based on the inferred user emotions.
[0234] [Postscript 8]
[0235] The system as described in Appendix 1 is characterized in that,
[0236] When analyzing video, the analysis unit detects specific actions or gestures and refines the analysis results accordingly.
[0237] [Postscript 9]
[0238] The system as described in Appendix 1 is characterized in that,
[0239] When analyzing video, the analysis unit analyzes background noise and ambient noise, and adds information about the working environment.
[0240] [Postscript 10]
[0241] The system as described in Appendix 1 is characterized in that,
[0242] The parsing unit infers the user's emotions and determines the priority of the parsing results based on the inferred user emotions.
[0243] [Postscript 11]
[0244] The system as described in Appendix 1 is characterized in that,
[0245] When parsing videos, the parsing unit refers to the user's past video parsing history to improve parsing accuracy.
[0246] [Postscript 12]
[0247] The system as described in Appendix 1 is characterized in that,
[0248] When parsing videos, the parsing unit considers the user's geographical location information to customize the parsing results.
[0249] [Postscript 13]
[0250] The system as described in Appendix 1 is characterized in that,
[0251] The generation unit estimates the user's emotions and adjusts the presentation of the generated images and descriptions based on the estimated user emotions.
[0252] [Postscript 14]
[0253] The system as described in Appendix 1 is characterized in that,
[0254] During the generation process, the generation unit generates images and explanatory text to emphasize specific scenes or important steps in the video.
[0255] [Postscript 15]
[0256] The system as described in Appendix 1 is characterized in that,
[0257] During the generation process, the generation unit generates explanatory text containing interactive elements based on the video content.
[0258] [Postscript 16]
[0259] The system as described in Appendix 1 is characterized in that,
[0260] The generation unit estimates the user's emotions and adjusts the length of the generated images and descriptions based on the estimated user emotions.
[0261] [Postscript 17]
[0262] The system as described in Appendix 1 is characterized in that,
[0263] When generating content, the generation unit refers to the user's past operation manual usage history to customize the generated content.
[0264] [Postscript 18]
[0265] The system as described in Appendix 1 is characterized in that,
[0266] The generation unit takes into account the user's device information during generation to generate optimal images and descriptions.
[0267] [Postscript 19]
[0268] The system as described in Appendix 1 is characterized in that,
[0269] The production department estimates the user's emotions and adjusts the structure of the operation manual based on the estimated user emotions.
[0270] [Postscript 20]
[0271] The system as described in Appendix 1 is characterized in that,
[0272] When creating the user manual, the production department refers to past user feedback to reflect areas for improvement.
[0273] [Postscript 21]
[0274] The system as described in Appendix 1 is characterized in that,
[0275] When creating the operation manual, the production department used a template customized based on the video content.
[0276] [Postscript 22]
[0277] The system as described in Appendix 1 is characterized in that,
[0278] The production department estimates the user's emotions and determines the priority of the operation manual based on the estimated user emotions.
[0279] [Postscript 23]
[0280] The system as described in Appendix 1 is characterized in that,
[0281] When creating the user manual, the production department takes into account the user's geographical location information to produce the optimal user manual.
[0282] [Postscript 24]
[0283] The system as described in Appendix 1 is characterized in that,
[0284] When creating the user manual, the production department analyzes users' social media activities and includes relevant information in the manual.
[0285] [Postscript 25]
[0286] The system as described in Appendix 1 is characterized in that,
[0287] The support unit infers the user's emotions and adjusts the chatbot's response method based on the inferred user emotions.
[0288] [Postscript 26]
[0289] The system as described in Appendix 1 is characterized in that,
[0290] The support department refers to the user's past question history to provide the best response when the chatbot answers.
[0291] [Postscript 27]
[0292] The system as described in Appendix 1 is characterized in that,
[0293] The support department provides customized responses based on the user's current situation or environment when the chatbot replies.
[0294] [Postscript 28]
[0295] The system as described in Appendix 1 is characterized in that,
[0296] The support unit estimates the user's emotions and determines the priority of the chatbot's responses based on the estimated user emotions.
[0297] [Postscript 29]
[0298] The system as described in Appendix 1 is characterized in that,
[0299] The support department considers the user's geographical location information to provide the best response when the chatbot replies.
[0300] [Postscript 30]
[0301] The system as described in Appendix 1 is characterized in that,
[0302] The support department analyzes users' social media activity and provides relevant information when the chatbot responds.
Claims
1. A system, characterized in that, include: The parsing unit is used to parse video. The generation unit is used to generate images and explanatory text based on the video content parsed by the parsing unit; The production department is used to create operation manuals based on the images and explanatory texts generated by the generation department. The support department is equipped with a chatbot for responding to operation manuals produced by the production department.
2. The system as described in claim 1, characterized in that, The analysis unit analyzes video content using generative AI.
3. The system as described in claim 1, characterized in that, The generation unit generates images and explanatory text for each step using generative AI.
4. The system as described in claim 1, characterized in that, The production department creates operation manuals based on images and explanatory texts generated by generative AI.
5. The system as described in claim 1, characterized in that, When users ask questions about the contents of the user manual, the support department uses a chatbot to answer them.
6. The system as described in claim 1, characterized in that, The support department uses generative AI for version management, ensuring that it always provides the latest user manuals.
7. The system as described in claim 1, characterized in that, The analysis unit infers the user's emotions and adjusts the video analysis method based on the inferred user emotions.
8. The system as described in claim 1, characterized in that, When analyzing video, the analysis unit detects specific actions or gestures and refines the analysis results accordingly.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A