Information processing system

By automatically acquiring and analyzing software source code, user manuals and FAQs are generated. Combined with image and video content, this solves the problem of information lag caused by manual writing in existing technologies, and realizes efficient and intelligent updating and optimization of user manuals and FAQs, thereby improving user experience and information management efficiency.

CN121900760APending Publication Date: 2026-04-21SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2025-10-16
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In the existing technology, the creation of user manuals and frequently asked questions (FAQs) for software products relies on manual labor, which is cumbersome and difficult to update in a timely manner. This results in inaccurate information for users, affecting the usability and user experience of the software, and makes it difficult to personalize and continuously optimize.

Method used

By automatically acquiring software source code, performing static analysis, extracting operational functions and error conditions, and using natural language generation technology to automatically generate user manuals and FAQs, combined with image and video content, automated and intelligent document updates and optimizations are achieved.

Benefits of technology

It enables efficient automatic generation and real-time updates of user manuals and FAQs, significantly reducing manual workload, improving user experience and information management efficiency, and adapting to software updates and user feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121900760A_ABST
    Figure CN121900760A_ABST
Patent Text Reader

Abstract

The present invention provides an information processing system comprising: means for acquiring a source code; the device is used for analyzing the acquired source code and extracting an operation function and an error condition; the natural language generation device is used for automatically generating a user manual and common questions and answers on the basis of the generated data; the rich media content generation device is used for automatically generating operation step pictures and videos; and the device is used for integrating the generated content into a management system and distributing the content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] In existing technologies, the creation and maintenance of user manuals and frequently asked questions (FAQs) for software products largely rely on manual labor. This process is cumbersome and easily lags behind software feature updates. When the source code changes, the relevant documentation cannot be updated in a timely manner, resulting in inaccurate information for users and affecting software usability and user experience. Furthermore, manually writing text and multimedia content is inefficient, makes personalization and continuous optimization difficult, and user feedback is often not reflected in the improvement of the documentation content in a timely manner. Summary of the Invention

[0004] To address the aforementioned problems, this invention provides an information processing system that automatically acquires software source code, performs static analysis to extract operational functions and error conditions, and automatically generates user manuals and FAQs using natural language generation technology. Simultaneously, the system can automatically generate rich media content such as operation step images and videos, and integrate all generated content into a management system for distribution. When changes to the source code are detected, the system automatically identifies the affected parts and updates the corresponding manuals and FAQs. The system also possesses user feedback collection and analysis capabilities, allowing for further optimization and improvement of related documents based on feedback information, achieving automated, intelligent, and continuous document updates, significantly reducing manual workload and enhancing user experience.

[0005] "Source code" refers to the computer program text used to describe the functions, structure, and behavior of a software system, including program instructions, comments, and data structures.

[0006] A "user manual" is an instructional document written to guide end users in the correct installation, operation, and maintenance of software products.

[0007] "Frequently Asked Questions (FAQ)" refers to a collection of documents that provide answers and explanations for frequently encountered questions by users while using the software.

[0008] "Natural language generation" refers to the process of automatically converting structured or semi-structured data into text information that conforms to the expression habits of natural language using artificial intelligence technology.

[0009] "Operational functions" refer to specific business processes or operation modules in a software system that users can access through the interface or API.

[0010] "Error conditions" refer to abnormal states, failure reasons, or errors caused by user misoperation that may occur during software operation.

[0011] "Rich media content" refers to content that includes diverse forms of expression such as images, audio, video, and animation, in addition to text information.

[0012] A "management system" refers to an information system that can centrally store, classify, manage, and intelligently distribute content such as documents, images, and videos.

[0013] "Distribution" refers to pushing the generated user manuals, FAQs, and rich media content to designated terminals, websites, or platforms via the network or other means for users to access.

[0014] "Feedback" refers to suggestions, opinions, or problem descriptions made by users after reading documents, manuals, or FAQs regarding the accuracy and usability of the content. Attached Figure Description

[0015] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0016] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0017] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0018] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0019] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0020] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.

[0021] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0022] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0023] Figure 9 This represents an emotion map that maps multiple emotions.

[0024] Figure 10 This represents an emotion map that maps multiple emotions.

[0025] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of Embodiment 1.

[0026] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0027] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system in Embodiment 2.

[0028] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0029] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0030] First, let me explain the terminology used in the following instructions.

[0031] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0032] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0033] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0034] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0035] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" is used to connect and express more than three items, the same interpretation as "A and / or B" applies.

[0036] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0037] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0038] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0039] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0040] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0041] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0042] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0043] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0044] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0045] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0046] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0047] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0048] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0049] Existing software product user manuals and frequently asked questions (FAQs) rely primarily on manual writing and maintenance, which is time-consuming and labor-intensive, and the information is difficult to update in a timely manner, resulting in users not receiving the latest and most accurate usage guidance and troubleshooting solutions. Furthermore, existing systems struggle to automatically and continuously improve documentation and related multimedia content based on code updates and user feedback, failing to effectively reduce the manpower costs of document management and updates, and also failing to adequately enhance the user support experience.

[0050] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0051] In this invention, the server includes a device for acquiring information, a device for statically parsing and processing information and automatically extracting user operation functions and abnormal states, a natural language content generation device for automatically constructing natural language input based on the extracted information and generating explanatory content through a generative artificial intelligence model, a multimedia information generation device for automatically generating corresponding visual information through image and video editing, and an electronic information management device for integrating, classifying, and distributing the generated content. This allows for the automatic acquisition and analysis of the latest information, and the efficient automatic generation, maintenance, and optimization of user manuals, FAQs, and multimedia explanatory content using artificial intelligence technology. This enables real-time updates and intelligent improvements to document information, significantly improving information management efficiency, reducing manual maintenance costs, and enhancing the end-user experience.

[0052] "Information" refers to a collection of data containing system function descriptions, operational logic, exception handling, etc., which can be electronic data in various forms such as source code, configuration files, and update logs.

[0053] "Static parsing processing" refers to the use of software tools to perform structured analysis on data or code without running the program, automatically identifying functional elements, anomaly judgment conditions, etc.

[0054] "User operation functions" refer to functional modules that users interact with on the terminal and trigger system responses, such as menu operations, button clicks, and data input.

[0055] "Abnormal state" refers to various abnormal working conditions that may occur during system operation, including errors, exceptions, exception prompts, fault information, etc.

[0056] "Natural language generation input information" refers to prompts or instruction texts used to drive generative artificial intelligence models to output natural language content, including functions to be described, operation procedures, and abnormal information.

[0057] "Generative AI models" refer to AI systems that are trained on large amounts of data and can automatically generate natural language text based on input, such as natural language processing models.

[0058] "Instructional content" refers to natural language text that describes functional modules, operating procedures, and handling of abnormalities, including user manuals, frequently asked questions, and other documents.

[0059] "Image editing and processing" refers to the editing, annotation, and compositing of images using software to generate visual explanatory materials.

[0060] "Video editing and processing" refers to using software to record, edit, and add voiceovers to operational procedures and other content to create dynamic demonstration and explanatory materials.

[0061] "Multimedia information" refers to composite information composed of various media forms such as text, images, and videos, which helps to improve the understanding and dissemination of content.

[0062] "Electronic information management device" refers to an electronic device or software system that can integrate, classify, store, and support the distribution of generated information.

[0063] "Distribution" refers to pushing generated content to user terminals or designated platforms through channels such as the internet for users to access and use.

[0064] "User feedback information" refers to the opinions, suggestions, questions, and other feedback information submitted by users regarding the generated content and system functions.

[0065] This invention relates to a system for automatically generating and updating user manuals and frequently asked questions (FAQs). A server, as the core device, enables the automatic acquisition, analysis, generation, and distribution of information. The system can run on hardware platforms such as general-purpose computer systems, cloud servers, and data processing equipment. Commonly used hardware includes a central processing unit (CPU), memory, and network interfaces. Software running on the server can include source code management tools (such as Git, SVN, etc.), static analysis tools (such as SonarQube, ESLint, etc.), generative artificial intelligence models (such as large-scale natural language processing models), image and video editing tools (such as Photoshop, Camtasia, etc.), and management systems for content storage and distribution (such as knowledge base platforms and document management software).

[0066] The server first interfaces with the code repository system to automatically retrieve the latest source code information for the required software. This process can be scheduled as a timed task to automatically fetch code via API, while the server records operation timestamps to ensure the information is up-to-date.

[0067] After acquiring the code, the server uses static analysis tools such as SonarQube and ESLint to analyze the source code. Through analysis, the server automatically identifies user operation functions (such as interface buttons, function entry points, and interface operations) and abnormal states (such as errors and exception messages). The extracted information is temporarily stored in structured data format for use in subsequent content generation.

[0068] Based on the information obtained from the analysis, the server automatically generates prompts for generative artificial intelligence models (such as natural language processing algorithms). For example, for the "Import Data" function, the server can generate the following prompt: Please provide step-by-step instructions for the 'Import Data' function, and list all possible errors and their solutions. The server inputs these prompts into a generative artificial intelligence model, automatically generating user manual content and FAQs in natural language. For example, the system generates the following FAQ: "The Smart Filter feature can help you quickly filter large amounts of data. Simply click 'Smart Filter' in the menu bar, enter the keywords you want to filter, such as 'age greater than 30', and the system will automatically highlight all data that meets the criteria. If the input criteria are entered in the wrong format, the system will prompt 'invalid filter criteria'. Please check whether the criteria format is correct, such as 'field name operator numeric'." In addition, the server will automatically or semi-automatically create multimedia content such as function flowcharts, user interface screenshots, and operation demonstration videos based on natural language content, using image editing tools (such as Photoshop) and video editing tools (such as Camtasia). The server annotates these multimedia files to facilitate user understanding.

[0069] After content generation is complete, the server integrates all text content and multimedia information, distributing it to terminals and users through knowledge base platforms, web portals, and other means. Users can access the latest manuals, FAQs, and supporting materials such as videos anytime via computers, mobile phones, and other terminals. If users encounter any problems during use, they can submit evaluations and feedback through the page. The server receives user evaluation information, automatically analyzes it, and drives subsequent optimization of the document content.

[0070] This invention reduces the cost of manual document writing and maintenance through a highly automated approach, enabling efficient updating and intelligent enhancement of product information, and effectively improving user experience and technical support efficiency.

[0071] use Figure 11 The processing procedure is explained.

[0072] Step 1: The server accesses the data management platform storing the source code and automatically pulls the latest source code of the target project. The input is the access address and credentials of the code repository, and the output is a new set of source code files stored locally or on the server. Specific operations include periodically triggering a code update script, downloading all or part of the necessary code after identity verification, and recording a timestamp for each pull operation.

[0073] Step 2: The server uses static analysis tools to analyze the acquired source code files. The input is the source code file obtained in step 1, and the output is the extracted functional descriptions and exception handling data structures. The server specifically performs the following actions: starting the static analysis task, traversing the source code file, identifying all functions, classes, call relationships, etc., and extracting explanatory text, comments, and conditional logic related to user operation functions and exception handling.

[0074] Step 3: Based on the static analysis results, the server automatically generates prompts for use by generative AI models. The input is the structured function and exception information output from step 2, and the output is AI-oriented natural language prompt text. The server combines each function and corresponding exception into a clear and explicit problem description or explanation of requirements, such as "Please explain the detailed operation steps of the 'Data Import' function, as well as possible errors and solutions."

[0075] Step 4: The server inputs the generated prompts into the generative artificial intelligence model to obtain the corresponding user manual instructions and FAQ text. The input is the natural language prompts from step 3, and the output is a detailed instruction document or FAQ entries. The server's specific actions include calling the AI ​​model via an interface, submitting each prompt to the model for processing, archiving and storing the text results returned by the AI, and setting necessary content tags and indexes.

[0076] Step 5: Based on the generated text content, the server uses image and video editing tools to create multimedia content such as screenshots of the user interface, diagrams, flowcharts, and operation demonstration videos. The input is the text content from step 4, and the output is corresponding image and video files. The server's actions include: automatically capturing screenshots of the actual software interface, adding captions to the images via scripts, recording or editing operation videos, and supplementing visual explanations.

[0077] Step 6: The server integrates all the organized text files, images, videos, and other multimedia content, and publishes them through a content management system or web portal. The input is all the content generated in step 5, and the output is a user-accessible document library or a page supporting various platforms. Server actions include: automatically generating a table of contents and index pages, setting user access permissions, and pushing notifications of the latest content changes to terminals or relevant users.

[0078] Step 7: Users access content published by the server through their terminals, and can browse knowledge bases, manuals, FAQs, and multimedia instructions. If they encounter problems, users can submit feedback. The input is the user's access request or feedback content, and the output is user evaluation data collected by the server. Specific user actions include browsing relevant content online, filling out feedback forms, and submitting them.

[0079] Step 8: The server analyzes the received user feedback and filters out frequently asked questions and suggestions. The input is the user-submitted feedback data, and the output is a list of content requiring optimization. The server then uses this data to return to step 2 or step 3, re-triggering the corresponding content generation and updates to achieve continuous automatic document optimization.

[0080] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0081] Existing user manuals and FAQs typically rely on manual writing and maintenance, resulting in content that fails to reflect the latest software or product status in real time, leaving users without timely and accurate information. Furthermore, the manuals cannot be personalized to the actual needs and emotional states of different users, making it difficult to provide timely guidance for resolving specific problems. In addition, user feedback is not effectively collected for continuous content optimization, leading to a lag in document updates, poor user experience, and low enterprise service efficiency. Therefore, there is an urgent need for an intelligent system that can automatically generate and optimize user manuals and FAQs, and dynamically adjust content based on user sentiment, to improve the timeliness, effectiveness, and satisfaction of users in obtaining information.

[0082] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0083] In this invention, the server includes a device for acquiring information data, a device for statically analyzing the information data to extract operational functions and abnormal states, a device for automatically generating natural language document data based on a generative artificial intelligence model, a device for automatically generating rich media content based on the natural language document data, a device for integrating and distributing the content to user devices, a device for detecting user emotional states and dynamically adjusting the distributed content, and a device for collecting user behavior and emotional data for content optimization. This enables the automatic generation and real-time updating of user manuals and frequently asked questions documents, personalized push notifications of auxiliary information based on user emotional states, and continuous optimization of content based on actual user usage and feedback, significantly improving user experience and information service efficiency.

[0084] "Information data" refers to raw or structured data that describes the functions, status, attributes, etc. of a target object, and may include source code, configuration files, product descriptions, user data, etc.

[0085] "Static analysis" refers to the process of analyzing and processing information data by checking its structure, semantics, and potential errors without executing any programs.

[0086] "Operational functions" refer to the specific operational behaviors or tasks that users can perform through the terminal, including various function calls, menu selections, and other operable items.

[0087] "Abnormal state" refers to an abnormal working mode that may occur during system operation or use, including errors, faults, abnormal branches, and abnormal prompts.

[0088] "Generative artificial intelligence models" refer to intelligent algorithm models that generate output content such as natural language, images, and audio by means of input prompts, based on deep learning or machine learning technologies.

[0089] "Natural Language Document Data" refers to descriptive document content written in natural language and output by generative artificial intelligence models, including user manuals, frequently asked questions, etc.

[0090] "Rich media content" refers to comprehensive display content that includes various information representation forms such as text, images, audio, and video, used to more intuitively assist users in understanding and operating the content.

[0091] "Information management device" refers to an information processing and publishing system used to store, integrate, manage, and distribute documents and rich media content.

[0092] "Wide area network" refers to a communication network system used to achieve data interconnection and information transmission between different physical locations, such as the Internet or enterprise private networks.

[0093] "Using device" refers to a terminal device used to access, display and interact with documents and rich media content, including computers, mobile terminals, tablets, etc.

[0094] "User emotional state" refers to the emotions and psychological state exhibited by users during use, including emotional characteristics such as confusion, satisfaction, and anxiety.

[0095] "Behavioral history" refers to all operation records and behavioral data generated when a user interacts with system content on a terminal device.

[0096] This invention can be implemented in the following specific ways.

[0097] This system comprises servers, terminals, and a wide area network, implemented using standard computing equipment and commonly used software tools. The servers can be high-performance computing devices running operating systems such as Linux, and equipped with automated processing scripts (such as Python), static analysis tools (such as SonarQube, pylint, cppcheck, etc.), generative artificial intelligence models (such as AI models based on deep learning platforms), rich media generation tools (such as DALL·E, FFmpeg, or open-source image / video processing software), and information management platforms (such as content management systems, CMS), among other related software. Terminal devices can be smartphones, tablets, or personal computers, equipped with display devices, cameras, microphones, and other input / output modules, and running applications or web pages that support data reception and display.

[0098] In the implementation of this invention, the server automatically acquires the source code, configuration files, or other information data of the product or software from the data source via a wide area network. The server uses static analysis tools to perform in-depth analysis of the acquired information data, automatically extracting the operational functions and abnormal states that users may encounter. The server constructs prompts from this structured data according to a preset template, and these prompts are sent as input to a generative artificial intelligence model. Based on the input content, the AI ​​model automatically generates natural language documents containing detailed operating instructions and frequently asked questions. The server can also call image generators, video editing tools, etc., to automatically generate relevant process images and operation demonstration videos, as well as other rich media content. All documents and rich media content are uniformly integrated into the information management system and distributed by the server to various terminal devices via the wide area network.

[0099] Users can access and browse user manuals, FAQs, and operation videos on their terminal devices. The terminal can automatically, or with user authorization, analyze and detect the user's emotional state during operation using its local camera, microphone, and emotion recognition algorithms (such as voice emotion APIs and facial expression recognition SDKs). The terminal transmits the collected emotional information back to the server in real time. Based on the user's emotional data and historical operation records, the server dynamically adjusts the content distributed to the terminal, such as automatically pushing more detailed step-by-step explanations, FAQ links, or supplementary videos to help users efficiently resolve practical difficulties. The server also regularly summarizes and analyzes all user interaction and emotional feedback data to continuously optimize the accuracy and relevance of AI-generated content, thereby achieving intelligent evolution and long-term self-improvement of the content.

[0100] For example, for a smart home appliance, the server can automatically obtain the device's function list and typical fault information through an IoT backend interface. For example: Product Name: Smart Cooking Machine Functions: Temperature adjustment, automatic cleaning Error modes: Heating failed, water tank low on water Requirement: Please generate a detailed operation guide and FAQ. Based on this, the server constructs the following prompt statement and sends it to the generative artificial intelligence model: Product Name: Smart Cooking Machine Description: Multifunctional cooking equipment Functions: Temperature adjustment, timer, automatic cleaning. Error modes: Heating failed, water tank low on water Please write a detailed user manual (including frequently asked questions) for the above products, and describe each function in simple and easy-to-understand natural language, and give examples of how to deal with errors. The AI ​​model automatically generates easy-to-understand manuals and FAQs for end users, which are then further processed by the server using image generation tools to create step-by-step operation illustrations or demonstration videos. If a user exhibits confusion or anxiety while browsing a chapter on the terminal, the server can automatically push "step-by-step explanation videos of the automatic cleaning function" or "common problem breakdowns" to help the user resolve the issue smoothly.

[0101] Through the above implementation methods, the present invention can automatically generate and continuously optimize product function descriptions and Q&A, and dynamically push content based on the user's individual emotional state, greatly improving user experience and service efficiency.

[0102] use Figure 12 The processing procedure is explained.

[0103] Step 1: The server automatically connects to the data source via a wide area network to obtain the latest software source code, configuration files, product descriptions, and other information. The input is raw data from a remote data warehouse, and the output is the raw data files downloaded to the server. The server executes pull commands (such as `git pull`) via scheduled scripts to ensure the most up-to-date data is obtained.

[0104] Step 2: The server uses static analysis tools (such as SonarQube or pylint) to perform in-depth analysis on the acquired information data. The input is locally stored source code or structured data, which the server parses, performs semantic analysis, and detects anomalies. The output is a structured list of operational functions and abnormal states (such as available functions, possible errors, etc.). The server automatically retrieves function definitions, error branches, and comment information from files to generate a feature data table.

[0105] Step 3: The server generates prompt statements conforming to a preset template based on structured operational functions and abnormal status data. The input consists of the function list and abnormal status set output from step 2, which are processed into natural language prompt statements. The output is text describing specific product functions and common errors. The server automatically assembles this content into complete sentences.

[0106] Step 4: The server invokes a generative artificial intelligence model, inputting the generated prompts into the model. It then automatically generates natural language documents such as user manuals and FAQs via API. The input is the prompts generated in step 3, and the output is detailed natural language operation instructions and frequently asked questions. The server automatically parses the AI ​​model output, formats the content, and archives it.

[0107] Step 5: The server uses image generation tools or video processing software to automatically create relevant operation flowcharts and instructional videos based on AI-generated document content. The input consists of operation steps described in natural language; the server decomposes the data and converts it into visual elements, flowcharts, or video scripts. The output consists of image and video files. The server generates and associates these files in batches with the corresponding operation sections.

[0108] Step 6: The server integrates all generated rich media content, including natural language documents, images, and videos, and uploads them uniformly to the information management system. The input consists of all the aforementioned content files, which the server organizes and archives based on content tags and metadata. The output is a structured content resource library accessible to end users. The server automatically synchronizes and updates.

[0109] Step 7: The server distributes the organized documents and rich media content to various terminals over the network. The input consists of relevant files from the content resource library. The server adaptively selects the push method (such as web pages, app push notifications, etc.) based on user and terminal information. The output is the content page that the user's terminal can access and display. The server monitors the distribution status and records user access behavior.

[0110] Step 8: The terminal displays user manuals, FAQs, operation images, and videos locally. Input is content files received from the server; the terminal automatically adjusts the display format based on screen size and performance. While operating the terminal, users can actively or passively trigger the camera and microphone to collect emotional data. Output is an interactive content page and the collected user behavior and emotional data.

[0111] Step 9: The terminal analyzes user behavior and emotional state, such as facial expressions and tone of voice, using a local emotion recognition algorithm. The input consists of audio and video data from the user's real-time interactions. The terminal intelligently analyzes this data and packages the emotion classification results (e.g., confusion, satisfaction, anxiety) along with the accessed content ID. The output is feedback data with emotion tags.

[0112] Step 10: The server receives user sentiment and behavioral feedback data from terminals and analyzes it in conjunction with user browsing history. Inputs include sentiment data packets uploaded by the terminal and historical interaction logs. The server determines the user's current state and filters corresponding supplementary content. Outputs include more targeted supplementary documents, FAQ links, or instructional videos. The server pushes supplementary resources in a targeted manner to improve user comprehension efficiency.

[0113] Step 11: Based on role learning and big data analysis, the server statistically analyzes all accumulated user behavior and emotional feedback. The input consists of a large volume of user operation history and sentiment tags; the server uses data mining tools to continuously optimize subsequent content generation patterns. The output is an automated content generation logic adapted and upgraded based on actual needs, continuously optimizing system service quality.

[0114] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0115] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0116] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0117] In this invention, the server includes a device for acquiring program information, a device for parsing program information and extracting functional information and abnormal condition information, a language information generation device for automatically generating natural language usage guidance information and question-and-answer information based on the extracted information, a visual information generation device for automatically generating multimedia visual information, an information management device for integrating and distributing information, and a control device for recognizing the user's emotional state based on the input device and dynamically adjusting the distributed content. This enables the system to automatically optimize content based on program changes and automatically adjust the information provided based on the user's emotional state and feedback, achieving efficient, personalized, and continuously optimized user assistance services and improving the overall user experience.

[0118] "Program information" refers to the collection of source code, scripts, configuration files, and related documents that describe the functions and operating logic of an application or software system.

[0119] "Functional information" refers to the list of software functions and their descriptions obtained by parsing program information, which are used for user operation.

[0120] "Abnormal condition information" refers to all state descriptions that may lead to unexpected behavior or errors, extracted from program information through static or dynamic analysis.

[0121] "Natural language" refers to spoken or written language used by humans in daily communication, such as Chinese and English, rather than computer programming languages.

[0122] "User guidance information" refers to documents, instructions, or text information that provide users with helpful content such as operating steps, methods, and precautions.

[0123] "Questions and answers information" refers to document content that provides answers, explanations, and suggestions for questions that users may ask, including a list of frequently asked questions.

[0124] "Language information generation device" refers to the hardware and software unit that automatically generates natural language text content, including modules that utilize generative artificial intelligence models.

[0125] "Multimedia information" refers to rich media content such as audio, images, animations, and videos used to help users understand the operation process.

[0126] "Visual information generation device" refers to a device or module that automatically generates multimedia content such as pictures, audio and video based on functional and operational information.

[0127] "Information management device" refers to a hardware and software system that integrates, classifies, stores, and distributes all generated text and multimedia information.

[0128] "Input device" refers to a hardware module that can collect input data such as user facial expressions, voice and operation behavior, such as a camera and microphone.

[0129] "User emotional state" refers to data or information collected and identified through input devices that represents a user's current psychological or emotional response.

[0130] "Control device" refers to the hardware and software unit that dynamically adjusts system behavior, manages content delivery, and personalizes configurations based on external detection results.

[0131] This invention can be implemented in the following ways.

[0132] The server connects to the data storage system via a network interface and first obtains the latest program information, such as source code, configuration files, or related scripts. In practice, the server can have an operating system installed, such as Linux, or it can run source code management software, such as code repository management tools. The server uses software such as version control tools to automatically retrieve and download the latest version of the program information based on the repository address and authentication information.

[0133] The server uses static analysis tools, such as static code analysis tools, to parse the acquired program information. Through the operation of the analysis tools, the server identifies and extracts various functional information in the system, as well as all possible abnormal condition information, such as program input and output, call logic, error handling, etc., and stores this data in a structured data format.

[0134] The server further invokes generative artificial intelligence models, such as general natural language processing models, to automatically generate user guidance and Q&A information. The server combines the extracted function list, exception condition information, and pre-defined prompts as input to the language model. For example, the server might use the following prompt: "Please write a user manual and frequently asked questions based on the following function list and error points." After the generative AI model returns the content, the server saves it as a user manual document or FAQ file for subsequent distribution and presentation.

[0135] The server can also utilize multimedia editing software, such as multimedia production tools, to automatically generate multimedia information, including operation images, videos, and animations, based on functional information and operation steps. The server can use automated scripts to generate various instructional videos according to the operation flow, or generate step-by-step diagrams using image editing technology.

[0136] The server is equipped with information management devices, such as a content management system, which integrates, classifies, and stores the generated text and multimedia information, and enables unified management and distribution to user terminals.

[0137] The terminal is equipped with input devices such as a camera and microphone. When users interact with or read content provided by the server, the terminal can use this hardware to collect the user's facial expressions, voice, and operational behaviors in real time. The terminal then uploads the collected data to an emotion recognition system via the network, such as a cloud-based or local emotion recognition model. After analyzing the data, the emotion recognition model outputs the user's current emotional state, such as confusion, satisfaction, or anxiety.

[0138] After receiving a user's emotional state, the server automatically adjusts the content pushed to that user. For example, if the server detects that a user is confused while browsing a particular instruction, it will automatically push more detailed instructions, operation examples, or tutorial videos. If the user shows satisfaction or understanding, the server reduces redundant information and only retains the core content. The server can also continuously optimize personalized content for subsequent users based on historical emotional states, usage behavior, and evaluation information.

[0139] For example, when a user is using a cooking app and needs to look up how to make a cake, if the user looks confused while watching the instructions, the terminal captures and uploads this expression. Once the server recognizes this state, it immediately pushes content such as "a detailed video of making a cake template" and "answers to common failure reasons." At this time, the server's prompt could be: "Please generate detailed step-by-step instructions and video demonstration suggestions for users searching 'how to make a cake' in the cooking app." This invention integrates multiple functions such as dynamic acquisition of program information, automatic analysis, AI content generation, multimedia production, emotion recognition, and dynamic content adjustment, to achieve a highly efficient information delivery system that is highly adaptable, timely in updating information, and supports personalized user needs.

[0140] use Figure 13 The processing procedure is explained.

[0141] Step 1: The server retrieves the latest program information from a remote repository via a network interface, including source code, configuration files, and script files. The server runs automated scripts that use tools such as version control to send requests and download relevant files. The output is a complete collection of program files stored locally.

[0142] Step 2: The server takes the program information from step 1 as input and calls static analysis tools (such as static code analysis tools) to parse the source code. The server automatically executes analysis commands to extract functional information and anomaly information. For example, the server parses out an API list, descriptions of each functional module, and potential error locations. The output is a structured data file, such as a list of functional and anomaly information.

[0143] Step 3: The server takes the extraction results from step 2 and the specified prompt as input, calls a generative artificial intelligence model to generate natural language content. The server constructs the input content, initiates an API call, and obtains the corresponding user guide text and Q&A content. The output is a user manual and FAQ document containing detailed explanations.

[0144] Step 4: The server takes the functional and operational steps from step 2 as input, calls multimedia generation and editing software, and automatically generates operational images and instructional videos. The server automatically triggers a media production script to convert the operational steps into images, videos, or animations. The output is a multimedia file, such as operational diagrams and demonstration videos.

[0145] Step 5: The server takes all the content generated in steps 3 and 4 as input and calls the content management module to organize, classify, and archive the information. The server writes text and multimedia information into a database or content management system according to metadata such as function and theme, facilitating subsequent distribution and searching. The output is a set of organized content accessible to the terminal.

[0146] Step 6: The terminal uses the user's actual actions and content browsing behavior as input, and uses its camera and microphone to collect the user's facial expressions and voice data. The terminal encrypts the collected raw data and sends it as input to the emotion recognition system. The output is the analyzed data of the user's current emotional state.

[0147] Step 7: The server takes the user's emotional state from step 6 as input, combines it with the distributable content in the content management system, and judges and adjusts the information to be pushed in real time. Based on different emotional states, the server automatically selects more suitable templates, detailed descriptions, or multimedia content, dynamically generates and pushes them to the terminal. The output is a personalized content page or multimedia resource.

[0148] Step 8: Users, as the primary consumers of content, take explanatory documents and multimedia content pushed by the server as input, and read, watch, or interact with them. If further questions arise, users can trigger a new round of interaction, with the terminal collecting new behaviors and emotions as new input, and the system iteratively optimizing its services. The output consists of log data on user satisfaction after receiving effective assistance and subsequent interactions.

[0149] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0150] Existing user manuals and frequently asked questions (FAQs) are typically statically generated, making it difficult to dynamically adjust them based on users' emotional states, comprehension abilities, and specific usage scenarios. This results in users not receiving timely personalized guidance and effective assistance when encountering problems during product use, thus reducing user experience and product satisfaction. Furthermore, existing systems struggle to respond promptly to program changes and user feedback, failing to automatically optimize guidance content, leading to untimely information updates and poor adaptability.

[0151] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0152] In this invention, the server includes a device for acquiring program description information, a device for parsing the program description information and extracting processing functions and abnormal conditions, a human language generation device for automatically generating information guidance text and question-and-answer sets based on the extracted data, a device for automatically generating multimedia content, a management device for integrating and distributing content, a device for recognizing user emotions using an information processing device and dynamically adjusting content based on a generative artificial intelligence model and prompts, and a device for displaying the adjusted content on a terminal. This enables the automatic generation and dynamic optimization of personalized operation guidance and multimedia auxiliary content based on the user's emotional state and real-time program changes, timely responding to user needs and improving information adaptability and user experience.

[0153] "Program description information" refers to a collection of text or data used to describe the structure, function, and operation of a computer program, including source code, scripts, or configuration files.

[0154] "Processing functions" refer to various functional operations or services implemented through programs, such as data processing and equipment control.

[0155] "Abnormal conditions" refer to abnormal states that may occur during program execution, including incorrect input, system failure, or abnormal operation.

[0156] "Information guidance text" refers to explanatory text content generated to guide users to operate correctly, including user manuals, operating procedures, etc.

[0157] A "question and answer set" refers to a collection of information that addresses common user questions and their corresponding answers.

[0158] "Human language generation device" refers to an information processing device that can automatically generate natural language text content based on input data.

[0159] "Multimedia content generation device" refers to a device that can automatically generate multimedia content such as pictures, audio, and video.

[0160] "Information management device" refers to a software or hardware system that integrates, manages, and distributes generated information content.

[0161] "Emotional information" refers to data that reflects a user's current emotional state, including facial expressions, tone of voice, and behavioral performance.

[0162] "Information processing device" refers to computing equipment used to collect, analyze and process various types of information (including emotional information).

[0163] "Generative artificial intelligence models" refer to artificial intelligence algorithm models based on technologies such as deep learning that can automatically generate text, multimedia, and other content based on input.

[0164] "Prompt statements" refer to instructional texts used to guide generative artificial intelligence models in generating the required content.

[0165] "Dynamically adjusting content" refers to optimizing and modifying existing content based on real-time information (such as user sentiment or program changes).

[0166] "Terminal" refers to a device used to display content and interact with users, such as a smartphone, tablet, or computer.

[0167] This invention relates to a system for dynamically generating and personally optimizing operation guidance and multimedia content based on a generative artificial intelligence model. The specific implementation of this invention will be described in detail below, referring to specific hardware and software names.

[0168] In this invention, the server can use a high-performance computer or cloud server as its hardware foundation, and the operating system can be Linux. The server first obtains program description information via the network, for example, by downloading the latest program source code from a code repository such as GitHub or GitLab using the git command. The server stores the source code locally and records its version information.

[0169] Subsequently, the server automatically parses the obtained program description information using static code analysis tools (such as SonarQube), identifying the defined processing functions (such as device operation instructions) and abnormal conditions (such as input errors, system errors, etc.). The server then converts the parsing results into structured data.

[0170] The server further utilizes generative artificial intelligence models (such as deep learning-based natural language processing algorithms, like GPT-3) to automatically generate operation guidance text and question-and-answer sets in conjunction with prompt statements. For example, the server can input the following prompt statement: "Please explain in a concise way how to operate the startMachine function. What should be done if the input is empty?" After receiving the text results returned by the AI ​​model, the server organizes this content into a user manual and frequently asked questions.

[0171] To further enrich the user experience, the server utilizes multimedia content generation tools (such as Stable Diffusion image generation tool and FFmpeg video processing software) to automatically create image descriptions and video demonstrations of the operation steps. This multimedia content, along with the text, is integrated into the information management device to form a complete information content set.

[0172] When a user accesses operation guides and FAQs through a terminal (such as a smartphone or tablet), the terminal uses its camera and microphone to capture the user's facial expressions and voice data. The terminal analyzes the user's emotional state through emotion recognition services (such as an emotion recognition system based on API calls). The terminal then feeds back the analysis results (such as "confused" or "satisfied") to the server in real time.

[0173] After receiving user emotion information, the server dynamically adjusts the content it provides based on this data, generative artificial intelligence models, and prompts. For example, when user confusion is detected, the server automatically generates more detailed video tutorials and instructions, and implements this through a prompt such as: "When the user shows signs of confusion, please generate a detailed step-by-step video tutorial." The terminal then presents personalized, dynamically optimized content to users in the form of text, images, and videos to help them resolve issues more quickly. Users can also provide feedback on the content through the terminal, such as clicking "Resolved" or "Having difficulty operating." The server continues to collect this feedback to further train and optimize the AI ​​model, continuously improving the relevance and adaptability of the content.

[0174] Specific examples: Suppose a user purchases a coffee machine and encounters an alarm and malfunctions during initial installation. The user consults the operating instructions via a mobile app but appears confused. The device's emotion recognition engine identifies the "confusion" emotion and sends it to the server. The server, using a generative AI model and the prompt "A detailed step-by-step guide to adding water to the coffee machine for beginners," automatically generates a personalized operating guide with step-by-step images and videos, which is then pushed to the user's device. The user successfully troubleshoots the problem using this guide and reports the result through their device. Based on this feedback, the server further optimizes subsequent content delivery.

[0175] Through the aforementioned devices and steps, this system achieves automated, personalized generation and dynamic optimization of operation instructions and multimedia content, effectively improving the user experience and problem-solving efficiency.

[0176] use Figure 14 The processing procedure is explained.

[0177] Step 1: The server retrieves the latest program description information (such as source code or configuration files) from the program code repository via a network interface. The input is the access address and authentication information of the remote repository, and the output is a locally stored program description information file. The server executes commands such as `git pull` to download the files and saves them in a specified local directory.

[0178] Step 2: The server uses static code analysis tools (such as SonarQube) to parse the obtained program description information. The input is a locally stored program description information file, and the output is structured data containing processing functions and exception conditions. The server calls the SonarQube API to upload and scan the code, and extracts the functional functions and possible error states from the scan results, organizing them into a specific data format.

[0179] Step 3: The server generates user manuals and Q&A sets based on structured data, using generative artificial intelligence models (such as natural language processing algorithms) and pre-designed prompts. Inputs include structured function and exception data and prompts; outputs are operation instructions and FAQ text expressed in natural language. The server sends content and prompts to the AI ​​model, such as "Please generate user operation instructions and exception handling methods for the startMachine function," and receives the text output.

[0180] Step 4: The server uses image generation tools (such as Stable Diffusion) and video editing tools (such as FFmpeg) to generate multimedia content. The input is a description of the operation steps, and the output is image and video files. The server automatically generates explanatory images for each step and automatically compiles video clips demonstrating the operation.

[0181] Step 5: The server integrates operational text and multimedia content, manages them uniformly through an information management system, and pushes the content to terminals. The input is all the generated content, and the output is a content package for the user terminal. The server automatically associates text, images, and videos, builds a content library, and prepares it for distribution.

[0182] Step 6: When a user accesses operation guides or FAQs, the terminal uses a camera and microphone to capture the user's facial expressions and voice data in real time. The input is the user's real-time audio and video data, and the output is raw emotional information data. The terminal calls upon hardware devices to record the data and uploads it to the emotion recognition service.

[0183] Step 7: The terminal analyzes the collected user information using emotion recognition algorithms (such as an emotion recognition API) to identify the user's current emotional state. The input is the collected audio and video data, and the output is the user's emotion tag (such as confusion or happiness). The terminal submits the uploaded data to the emotion API and receives the analysis results.

[0184] Step 8: The server receives sentiment analysis results and, based on the user's current emotional state and the content context, dynamically generates or adjusts operation instructions and multimedia content using a generative artificial intelligence model and prompts. Inputs include sentiment tags, contextual information, and prompts; outputs are personalized, adjusted content. For example, the server might input "The user appears confused; please generate a more detailed step-by-step video tutorial," and the AI ​​model will return personalized content.

[0185] Step 9: The terminal receives personalized, dynamically adjusted instruction text and multimedia content pushed by the server and displays it through a video player or application interface. The input is the latest content package, and the output is user-visualized and actionable guidance content. The terminal automatically pops up videos, displays images and text to facilitate user understanding and operation.

[0186] Step 10: Users interact with products based on the content displayed on the terminal and submit feedback through the terminal. Input consists of the user's actual actions and feedback, while output is the user feedback data. After completing the operation, users can click "This content has resolved the issue" or submit "Unresolved" feedback to help the server optimize subsequent content generation and service strategies.

[0187] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0188] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0189] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0190] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0191] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0192] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0193] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0194] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0195] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0196] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0197] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0198] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0199] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0200] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0201] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0202] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0203] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0204] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0205] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0206] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0207] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0208] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0209] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0210] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0211] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0212] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0213] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0214] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0215] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0216] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0217] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0218] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0219] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0220] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0221] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0222] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0223] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0224] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0225] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0226] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0227] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0228] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0229] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0230] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, as well as inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0231] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0232] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0233] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0234] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0235] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0236] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0237] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0238] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0239] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0240] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0241] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0242] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0243] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0244] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0245] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0246] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0247] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0248] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0249] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0250] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0251] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0252] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0253] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0254] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0255] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0256] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0257] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0258] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0259] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0260] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0261] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0262] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0263] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0264] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0265] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0266] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0267] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0268] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0269] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0270] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0271] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0272] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0273] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0274] In addition, the following notes are provided in response to the above explanation.

[0275] Example 1 (Note 1) An information processing system includes: a device for acquiring information; a device for processing the acquired information through static parsing and automatically extracting user operation functions and abnormal states; a natural language content generation device for automatically constructing the input information required for natural language generation based on the extracted information and generating explanatory content through a generative artificial intelligence model; a multimedia information generation device for automatically generating multiple visual information through image editing and video editing based on the generated explanatory content; and a device for integrating, classifying, and distributing the generated information through an electronic information management device.

[0276] (Note 2) The information processing system according to Appendix 1 further includes a processing device for detecting changes in information content and automatically updating natural language content and multimedia information based on the changed portions.

[0277] (Note 3) The information processing system according to Appendix 1 further includes a processing device for acquiring user evaluation information and automatically improving natural language content and multimedia information based on the evaluation information.

[0278] Application Example 1 (Note 1) An information processing system includes: a device for acquiring information data; a device for performing static analysis on the acquired information data to extract operational functions and abnormal states; a device for automatically generating natural language document data using a generative artificial intelligence model based on the extracted operational functions and abnormal states; a device for automatically generating rich media content containing visual and audio information based on the natural language document data; a device for integrating the generated natural language document data and rich media content into an information management device and distributing it to a user device via a wide area communication network; a device for detecting the user's emotional state when the user device displays the natural language document data and rich media content and sending the detection result to the information processing device; a device for dynamically adjusting the distributed information content according to the detected user emotional state; and a device for collecting user behavior history and emotional state data and using it to optimize the content.

[0279] (Note 2) The information processing system described in Appendix 1 can automatically detect updates or changes in information data, and based on the detection results, automatically regenerate natural language document data and rich media content using a generative artificial intelligence model and distribute them.

[0280] (Note 3) The information processing system described in Appendix 1 is capable of analyzing user usage status information and emotional status information, and automatically optimizing the content of natural language document data and rich media content based on the analysis results.

[0281] Example 2 (Note 1) An information processing system includes: a device for acquiring program information; a device for parsing the acquired program information and extracting functional information and abnormal condition information; a language information generation device for automatically generating natural language usage guidance information and question-and-answer information based on the extracted information; a visual information generation device for automatically generating multimedia information based on functional information and operation information; an information management device for integrating the generated information into an information management device and distributing it; and a control device for recognizing the user's emotional state through an input device and dynamically adjusting the distributed information according to the recognition result.

[0282] (Note 2) The information processing system according to Appendix 1 further includes a control device for detecting changes in program information and automatically updating usage guidance information and question-and-answer information based on the detected changes.

[0283] (Note 3) The information processing system according to Appendix 1 further includes a control device for parsing usage history information and evaluation information obtained from users, and reflecting the parsed results in the improvement of usage guidance information and Q&A information.

[0284] Application Example 2 (Note 1) An information processing system includes: a device for acquiring program description information; a device for parsing the acquired program description information and extracting processing functions and abnormal conditions; a human language generation device for automatically generating information guidance text and question-and-answer sets based on the extracted data; a multimedia content generation device for automatically generating image and video information about operation steps; a device for integrating the generated content into an information management device and distributing it; a device for dynamically adjusting the provided content based on the recognition results using a generative artificial intelligence model and prompt statements, utilizing an information processing device capable of recognizing emotional information; and a device for displaying the adjusted content on a user terminal.

[0285] (Note 2) The information processing system according to Appendix 1 further includes a device for detecting changes in program description information and automatically updating information guidance text and question-and-answer set according to the changes.

[0286] (Note 3) The information processing system according to Appendix 1 further includes means for acquiring user feedback information and improving information guidance text and question-and-answer set content based on the feedback information.

Claims

1. An information processing system, characterized in that, include: Device for obtaining source code; A device for parsing the acquired source code and extracting operational functions and error conditions; A natural language generation device for automatically generating user manuals and frequently asked questions based on the generated data; A rich media content generation device for automatically generating images and videos of operation steps; and A device for integrating and distributing generated content into a management system.

2. The information processing system according to claim 1, characterized in that, It also includes a device that detects source code changes and automatically updates the user manual and FAQs based on the changes.

3. The information processing system according to claim 1, characterized in that, It also includes means for receiving user feedback and analyzing the feedback to improve the user manual and frequently asked questions.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A