system
A system that collects, preprocesses, and trains CEO-related data to generate a virtual CEO emulator with AI, addressing the challenge of CEO time constraints and enhancing corporate efficiency and decision-making accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-02
- Publication Date
- 2026-04-14
AI Technical Summary
In modern business environments, CEOs face challenges in being directly involved in all scenarios due to time and resource limitations, and there are limited means to inherit and utilize their knowledge and judgment over time, necessitating a mechanism to mimic their knowledge and judgment in real time.
A system that collects CEO-related information, preprocesses it to remove noise and standardize formats, trains a generative AI model, and interacts through a user interface, continuously retraining the model to generate a virtual CEO emulator with speech recognition and synthesis capabilities.
Enables the utilization of a CEO's knowledge and judgment in real time, improving corporate operations' efficiency and decision-making accuracy.
Smart Images

Figure 2026064610000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the modern business environment, the decision-making and statements of the CEO are very important for corporate operations. However, due to physically limited time and resources, it is difficult for the CEO to be directly involved in all scenarios. Furthermore, there are limited means to inherit and utilize the CEO's knowledge and judgment over time. To solve these problems, there is a need for a mechanism to mimic the CEO's knowledge and judgment in real time and generate a virtual CEO that lives forever in two dimensions.
Means for Solving the Problems
[0005] To solve the above-mentioned problems, the present invention proposes the following technical means. First, it provides means for collecting all information related to the CEO. Next, it provides means for preprocessing the collected information, including noise reduction and standardization of data formats. It also provides means for training a generative AI model using the preprocessed data to generate a virtual CEO emulator. Furthermore, it proposes constructing a system that includes means for interacting with the virtual CEO through a user interface, and means for collecting interaction data and continuously retraining the AI model. In addition, it provides means for analyzing the preprocessed data using natural language processing technology to extract important keywords and sentences, and means for generating a virtual CEO emulator as a two-dimensional interface and interacting with the user through speech recognition and speech synthesis functions. This realizes a system that can utilize the CEO's knowledge and judgment for the future.
[0006] "CEO" refers to the chief executive officer of a company, and is responsible for making strategic decisions for the entire company.
[0007] "Information" refers to data such as statements, decisions, emails, social media posts, and interview articles related to the CEO.
[0008] "Data collection" refers to the process of gathering various types of data through web scraping techniques or manual input from users.
[0009] "Preprocessing" refers to all processes that remove noise from collected data and standardize the data format.
[0010] "Noise" refers to information that is unnecessary or irrelevant during data analysis or model training.
[0011] "Data format standardization" refers to the process of converting data into a specific, parseable format.
[0012] A "generative AI model" refers to an artificial intelligence model that generates new text or dialogue based on input data.
[0013] A "virtual CEO emulator" refers to a system that uses generative AI models to mimic the statements and decisions of a CEO.
[0014] "User interface" refers to an interactive screen or platform through which a user interacts with a virtual CEO emulator.
[0015] "Dialogue" refers to the communication exchange that takes place between the user and the virtual CEO emulator.
[0016] "Interaction data" refers to the content of the conversations that take place between the user and the virtual CEO emulator.
[0017] "Retraining" refers to the process of updating a generative AI model using new data to improve its performance.
[0018] "Natural language processing technology" refers to the technology used to analyze and understand human language using computers.
[0019] A "keyword" refers to a word that is considered particularly important within text data.
[0020] A "sentence" refers to a word or text, and is a fundamental unit of analysis in natural language processing.
[0021] A "2D interface" refers to a graphical user interface designed to allow the user to visually interact with a virtual CEO emulator.
[0022] "Speech recognition" refers to the technology that analyzes voice input and converts it into text data.
[0023] "Speech synthesis" refers to the technology that generates human speech based on text data. [Brief explanation of the drawing]
[0024] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0025] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0026] First, let's explain the terminology used in the following explanation.
[0027] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).
[0028] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0029] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0030] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0031] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0032] [First Embodiment]
[0033] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0034] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0035] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0036] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0037] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0038] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0039] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0040] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0041] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0042] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0043] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0044] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0045] System Overview
[0046] This invention relates to a system that generates a virtual CEO emulator by collecting and preprocessing information related to the CEO and training a generative AI model, enabling interaction through a user interface. This system can learn the CEO's statements and decisions, imitate them in real time, and interact with the user via a two-dimensional interface.
[0047] Program Processing Overview
[0048] This system consists of the following main phases:
[0049] 1. Data Collection Phase
[0050] 2. Data preprocessing phase
[0051] 3. Model training phase
[0052] 4. CEO Emulator Generation Phase
[0053] 5. Interaction Phase
[0054] Program operation description
[0055] Data collection phase
[0056] The server collects publicly available information related to the CEO using web scraping techniques. This includes social media posts, blog articles, interviews, and publicly available emails.
[0057] Users can improve the accuracy of data collection by uploading internal documents and files to the system.
[0058] Data preprocessing phase
[0059] The server preprocesses the collected data, removing noise and standardizing the data format. Specifically, it uses a grammar checking tool to correct typos and removes unnecessary HTML tags and special characters.
[0060] The server uses natural language processing technology to analyze text data and extract important keywords and sentences.
[0061] Model Learning Phase
[0062] The server uses a generative AI model (e.g., GPT-3® or BERT) to train on preprocessed data. Initial setup is performed, and the model is instantiated.
[0063] The server splits the data into a training set and a test set and performs training. During this process, it optimizes the model parameters using backpropagation.
[0064] The server evaluates the model's performance using a test set and retrains it as needed.
[0065] CEO Emulator Generation Phase
[0066] The server generates a virtual CEO emulator as a 2D interface. This emulator is equipped with speech recognition and speech synthesis capabilities, enabling real-time interaction with the user.
[0067] The server deploys the generated interface as a user interface on the cloud, making it accessible to users.
[0068] Interaction Phase
[0069] Users interact with a virtual CEO through their device. For example, they log in to a web application using a browser and open a dialogue screen.
[0070] The server analyzes user input (text or voice) and generates an appropriate response in real time. The generated response is then delivered to the user as voice output through a speech synthesis engine.
[0071] The server continuously collects new dialogue data and retrains the model, thereby improving the system's accuracy and response quality.
[0072] Specific example
[0073] Data collection examples
[0074] The server collects the CEO's social media posts and interview articles from the past five years and stores them in a database.
[0075] Users can supplement the collected data by uploading personal emails and internal documents to the system.
[0076] Data preprocessing example
[0077] The server analyzes the collected data and corrects typos and grammatical errors. It also removes unnecessary tags and special characters, converting the data into clean text.
[0078] Model Learning Example
[0079] The server uses pre-processed data to train a generative AI model, learning the CEO's speaking patterns. For example, it incorporates frequently used phrases and decision-making tendencies into the model.
[0080] Emulator generation example
[0081] The server generates a 2D virtual CEO and implements speech recognition and synthesis capabilities. These emulators are deployed on the cloud as the user interface.
[0082] User interaction examples
[0083] Users ask questions about project progress to a virtual CEO and receive specific advice. The server analyzes the user's questions in real time and generates the most appropriate answers.
[0084] In summary, this invention provides a system that allows CEOs to leverage their knowledge and judgment for the future. This will lead to increased efficiency in corporate operations and improved accuracy in decision-making.
[0085] The following describes the processing flow.
[0086] Step 1: Data Collection
[0087] The server uses web scraping techniques to collect data related to the CEO from specified websites and social media platforms. This includes publicly available blog posts, interviews, and social media posts by the CEO.
[0088] Users can use the system interface to upload additional CEO-related data, such as internal documents and past email data.
[0089] Step 2: Save
[0090] The server stores the collected data in a database for centralized management. The stored data includes various formats such as text, images, and audio files.
[0091] Step 3: Noise Reduction and Data Cleansing
[0092] The server performs noise reduction processing on the collected data. This includes tasks such as correcting spelling errors and removing unnecessary tags and special characters.
[0093] The server standardizes data formats and cleans up text data. For example, it removes HTML tags and converts it to plain text.
[0094] Step 4: Data analysis using natural language processing
[0095] The server uses natural language processing techniques to analyze clean text data. It performs tokenization, part-of-speech tagging, and sentence analysis to extract important keywords and sentences.
[0096] Step 5: Create a dataset
[0097] The server splits the analyzed data into a training set and a test set. For example, 70% of the data might be used for the training set and 30% for the test set.
[0098] Step 6: Training the Generative AI Model
[0099] The server trains a generative AI model using a training set. It uses deep learning algorithms and performs backpropagation to optimize the model's parameters.
[0100] The server trains the model multiple times, specifying the number of epochs, to improve the model's accuracy.
[0101] Step 7: Model Evaluation
[0102] The server evaluates the performance of the trained model using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[0103] The server retrains the model as needed to further improve its performance.
[0104] Step 8: Generating the CEO Emulator
[0105] The server uses the final generative AI model to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time.
[0106] Step 9: Deploying the User Interface
[0107] The server deploys the generated virtual CEO emulator to the cloud. It provides a URL that users can access via a browser, and functions as the user interface.
[0108] Step 10: Start interacting with the user
[0109] Users initiate a conversation with a virtual CEO using a web interface. They can ask questions and seek advice using text or voice input.
[0110] The server analyzes user input in real time and generates an appropriate response. The generated response is then delivered to the user as speech output using a speech synthesis engine.
[0111] Step 11: Collect interaction data and retrain
[0112] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement.
[0113] The server uses the collected interaction data to retrain the model and improve the quality of the dialogue.
[0114] (Example 1)
[0115] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0116] In today's business environment, there is a growing need to quickly and accurately mimic the judgments and instructions of managers in order to improve the efficiency of corporate operations and the precision of decision-making. Furthermore, there is an increasing need for systems that can utilize the knowledge and experience of managers even when they are physically absent. However, conventional technologies do not adequately provide methods for mimicking managers' thought patterns and statements in real time.
[0117] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0118] In this invention, the server includes means for collecting publicly available information, means for preprocessing the collected information, performing grammar checks and unifying data formats, means for training a generative AI model using the preprocessed data to generate a virtual manager emulator, means for interacting with the virtual manager through a user interface, and means for collecting interaction data and continuously retraining the AI model. This enables real-time imitation of the manager's knowledge and judgment, improving the efficiency of business operations and the accuracy of decision-making.
[0119] "Public information" refers to information that is accessible to a large number of people, such as information posted on websites, social media posts, blog articles, interview articles, and publicly available emails.
[0120] "Preprocessing" is the process of correcting typographical errors in collected data, removing unnecessary HTML tags and special characters, and standardizing data formats.
[0121] A "generative AI model" is an artificial intelligence model that is trained using collected and pre-processed data, and includes, for example, natural language generation models such as GPT-3.
[0122] A "virtual executive emulator" is a virtual character that mimics the statements and decisions of executives in real time based on a generated AI model, and it has speech recognition and speech synthesis capabilities.
[0123] "User interface" refers to the screens and applications that allow a user to interact with a virtual manager emulator, and includes web browsers and dedicated applications.
[0124] "Interaction data" refers to the data from conversations between the user and a virtual manager emulator, and this data is used to retrain the AI model.
[0125] "Natural language processing technology" refers to techniques for analyzing text data and extracting important keywords and sentences, and includes libraries such as NLTK and SpaCy.
[0126] "Speech recognition" is a technology that converts speech data into text data, and an example of this is Google® Speech-to-Text.
[0127] "Speech synthesis" is a technology that converts text data into speech data, and Amazon Polly is an example of this.
[0128] This invention is a system that generates a virtual CEO emulator by collecting and pre-processing information related to the CEO and training a generated AI model, enabling interaction through a user interface. This system imitates the judgments and statements of executives in real time based on publicly available information, thereby improving the efficiency of corporate operations and the accuracy of decision-making.
[0129] The system configuration is mainly described in the following phases:
[0130] 1. Data Collection Phase
[0131] 2. Data preprocessing phase
[0132] 3. Model training phase
[0133] 4. CEO Emulator Generation Phase
[0134] 5. Interaction Phase
[0135] Data collection phase
[0136] The server uses web scraping techniques (e.g., Beautiful Soup) to collect publicly available information such as social media posts, blog articles, interviews, and published emails. Specifically, it uses Python libraries to crawl websites containing certain keywords (e.g., names of business owners, company names).
[0137] Users can improve the quality and quantity of data collection by uploading internal documents and personal emails to the system. The uploaded files are stored in a database on the server.
[0138] Data preprocessing phase
[0139] The server corrects the collected data using a grammar checking tool (e.g., Grammarly API) to eliminate typos and grammatical errors.
[0140] The server removes unnecessary HTML tags and special characters from the collected data, converting it into clean text data.
[0141] The server analyzes text data using natural language processing technologies (e.g., NLTK, SpaCy) to extract important keywords and sentences. For example, it can extract key topics and keywords from interview articles with business executives.
[0142] Model Learning Phase
[0143] The server instantiates a generative AI model (e.g., GPT-3) and loads pre-processed data. Instances of the generative AI model are created using the OpenAI® API.
[0144] The server splits the preprocessed data into a training set and a test set, and then trains the model.
[0145] The server uses backpropagation to optimize the parameters of the generated AI model and retrains it to improve its accuracy.
[0146] CEO Emulator Generation Phase
[0147] The server generates a virtual manager emulator using a pre-trained generative AI model. This emulator is designed as a two-dimensional interface and includes speech recognition (e.g., Google Speech-to-Text) and speech synthesis (e.g., Amazon Polly) capabilities.
[0148] The server deploys a virtual manager emulator on the cloud and makes it accessible to users. For example, it might be deployed on AWS® and configured to be accessible to users via a browser.
[0149] Interaction Phase
[0150] The user logs into the web application from their device and opens the interactive interface.
[0151] Example: A user logs into the system using a browser and asks a virtual manager, "What is the current project status?"
[0152] The server analyzes user input (text or voice) and generates an appropriate response.
[0153] Example: The server analyzes the user's question and provides the generated response as speech output using a speech synthesis engine.
[0154] The server continuously collects new interaction data and retrains the generative AI model. This improves the system's accuracy and response quality.
[0155] Examples of specific prompt statements include the following:
[0156] "Could you tell me the CEO's view on the progress of the current project?"
[0157] "We'd like to hear the CEO's opinion on our new marketing strategy."
[0158] As described above, the present invention makes it possible to mimic the knowledge and judgment of managers in real time, thereby improving the efficiency of corporate operations and the accuracy of decision-making.
[0159] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0160] The specific processing flow of the program
[0161] Step 1: Data Collection
[0162] The server launches a web scraping script to collect publicly available information such as social media posts, blog articles, interviews, and published emails. Specifically, it uses the Python Beautiful Soup library to crawl websites containing specific keywords.
[0163] Input: Website URL or specific keywords
[0164] Output: Collected text data (e.g., text from social media posts or blog articles)
[0165] Specific operation: The server accesses a specific website, retrieves the page content, and extracts the necessary parts.
[0166] Step 2: Data Upload
[0167] Users upload internal documents and personal emails to the system. The uploaded files are stored in a database on the server.
[0168] Input: A file selected by the user (e.g., PDF, text file)
[0169] Output: Text data stored in the database
[0170] Specific operation: The user sends a file to the server using the system's upload form, and the server saves its contents to the database.
[0171] Step 3: Data Cleaning
[0172] The server uses a grammar checking tool (e.g., Grammarly API) to correct typos and grammatical errors in the collected data.
[0173] Input: Collected text data
[0174] Output: Clean text data with typos and grammatical errors corrected.
[0175] Specific operation: The server sends text data to a grammar checking tool and receives the corrected text.
[0176] Step 4: Unify text formatting
[0177] The server removes unnecessary HTML tags and special characters from the collected data to standardize the data format.
[0178] Input: Text data with typos and grammatical errors corrected.
[0179] Output: Clean text data
[0180] Specific operation: The server uses regular expressions to remove HTML tags and special characters, converting the data into clean text.
[0181] Step 5: Natural Language Processing (NLP)
[0182] The server uses natural language processing techniques (e.g., NLTK, SpaCy) to analyze text data and extract important keywords and sentences.
[0183] Input: Clean text data
[0184] Output: Extracted keywords and important sentences
[0185] Specific operation: The server starts a natural language processing library, analyzes the text data, and extracts important information.
[0186] Step 6: Model instantiation and data loading
[0187] The server instantiates a generative AI model (e.g., GPT-3) and loads pre-processed data.
[0188] Input: Extracted keywords and important sentences
[0189] Output: Text data loaded as training data
[0190] Specific operation: The server uses the OpenAI API to create a GPT-3 instance and load the pre-processed data.
[0191] Step 7: Model Training
[0192] The server divides the pre-processed data into a training set and a test set, and then trains the generative AI model.
[0193] Input: Loaded training data
[0194] Output: Trained generative AI model
[0195] Specific operation: The server optimizes the model using training data and evaluates its accuracy using test data.
[0196] Step 8: Generate CEO emulator
[0197] The server generates a virtual manager emulator using a pre-trained generative AI model. This emulator has a two-dimensional interface and features speech recognition and speech synthesis capabilities.
[0198] Input: Trained generative AI model
[0199] Output: Virtual manager emulator
[0200] Specific operation: The server generates characters from the model and integrates speech recognition and speech synthesis engines.
[0201] Step 9: Interface Deployment
[0202] The server deploys a virtual manager emulator on the cloud, making it accessible to users.
[0203] Input: Virtual manager emulator
[0204] Output: Deployed user interface
[0205] Specific operation: The server deploys the emulator to a cloud service such as AWS and configures it so that users can access it via a browser.
[0206] Step 10: User Interaction
[0207] The user logs into the web application from their device, opens a conversational interface, and interacts with a virtual manager.
[0208] Input: User voice or text input
[0209] Output: Response from the virtual manager
[0210] Specific operation: The user inputs a question or instruction, and the server generates a response and sends it back.
[0211] Step 11: Data Collection and Retraining
[0212] The server continuously collects conversational data between the user and the virtual manager, and uses it to retrain the generated AI model.
[0213] Input: User interaction data
[0214] Output: Retrained generative AI model
[0215] Specific operation: The server collects interaction data, periodically retrains the model, and improves the system's accuracy and response quality.
[0216] Through the processing steps described above, this system can mimic the knowledge and judgment of managers in real time, thereby improving the efficiency of business operations and the accuracy of decision-making.
[0217] (Application Example 1)
[0218] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0219] In recent years, there has been a growing demand for increased efficiency in corporate management and improved accuracy in decision-making. However, means of passing on leaders' knowledge and judgment have been limited, and providing real-time advice has been difficult. In particular, there is a need for improved quality of immediate information provision and advice to users in virtual environments. Furthermore, when using virtual reality devices, the lack of appropriate conversational interfaces to enhance the user experience is a challenge.
[0220] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0221] In this invention, the server includes means for collecting information, means for preprocessing the collected information to remove noise and unify the data format, means for training a generative AI model using the preprocessed data to generate a virtual mentor emulator, means for interacting with the virtual mentor through a user interface, means for collecting interaction data and continuously retraining the AI model, and means for providing the user with advice on product selection within a virtual environment through a virtual reality device. This enables increased efficiency in business management and improved accuracy in decision-making, and provides a real-time conversational interface that enhances the user experience in a virtual environment.
[0222] "Information" refers to a series of data related to the user, such as documents, data, audio, and images.
[0223] "Collection" refers to the entire process of gathering information.
[0224] "Preprocessing" refers to the process of removing noise from collected information and standardizing the data format.
[0225] A "generative AI model" refers to an algorithm or system that uses a large dataset for artificial intelligence to learn and perform a specified task.
[0226] A "virtual leader emulator" refers to a virtual agent created to mimic the speech patterns and behavioral patterns of a specific leader.
[0227] A "user interface" refers to the interface through which a user interacts with a system.
[0228] "Interaction data" refers to data related to the interactions and operations that take place between the user and the system.
[0229] "Retraining an AI model" refers to the process of retraining an existing AI model to improve its performance based on new data.
[0230] "Virtual reality devices" refer to hardware and software used to provide a virtual reality (VR) environment.
[0231] "Product selection advice" refers to information and suggestions provided by a supervisor emulator to help users choose a specific product.
[0232] "Real-time" refers to the process of generating an immediate response to user input.
[0233] System Overview
[0234] This invention realizes a system that generates a virtual instructor emulator and provides advice to users through a virtual environment or virtual reality device. Specifically, users can move freely within a virtual store and receive real-time advice from a virtual instructor regarding product selection.
[0235] Hardware and software used
[0236] Server: Performs data collection, preprocessing, training and retraining of AI models, and generation of virtual leader emulators.
[0237] Virtual reality devices (HMDs, etc.): Used by users to move around within a virtual store and interact with a virtual instructor emulator through an interface.
[0238] Generative AI Models: Using GPT-3 or similar high-performance AI models, the system generates statements and advice from a virtual leader emulator.
[0239] Natural language processing tools: Used for analyzing and preprocessing collected text data.
[0240] Processing flow and data calculations
[0241] 1. Information Gathering: The server uses web scraping techniques to collect publicly available information and internal documents. This includes social media posts, interview articles, and publicly available emails.
[0242] 2. Preprocessing: The collected information is preprocessed to remove noise and standardize the data format. Specifically, grammar checking tools are used to correct typographical errors and unnecessary HTML tags and special characters are removed.
[0243] 3. Training Generative AI Models: GPT-3 or similar generative AI models are trained using preprocessed data. During training, the data is divided into a training set and a test set, and the model is optimized.
[0244] 4. Generation of a virtual leader emulator: The server generates a virtual leader emulator using a trained AI model. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time.
[0245] 5. Collection and retraining of interaction data: Data generated when users interact with the instructor emulator using a virtual reality device is collected and used to continuously retrain the AI model.
[0246] Specific example
[0247] In a virtual store:
[0248] The user wears an HMD (Head-Mounted Display) and moves around within the virtual store.
[0249] The participant makes the gesture of picking up any product and is asked, "What are the features of this product?"
[0250] A virtual coach emulator provides real-time responses such as, "This product utilizes the latest running technology, is extremely lightweight, and durable. It is especially suitable for long-distance running."
[0251] Example of a prompt
[0252] User: "Could you please give me some advice about this new smartphone?"
[0253] Virtual Leader Emulator: "This smartphone features the latest processor, making it extremely fast and offering excellent battery life. It's especially recommended for gaming and video editing."
[0254] Thus, this invention allows users to receive real-time advice from a mentor emulator when effectively selecting products in a virtual environment. This aims to improve the efficiency of business operations and the accuracy of decision-making.
[0255] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0256] Step 1:
[0257] The server collects information. Specifically, it uses web scraping techniques to collect publicly available information from the web (such as social media posts, blog articles, interviews, and publicly available emails) and stores it in a database. The input is a URL or API endpoint on the web, and the output is the collected text data.
[0258] Step 2:
[0259] The server preprocesses the collected information. Specifically, it uses a grammar checker to correct typos and removes HTML tags and special characters. Next, it analyzes the text data using natural language processing techniques to extract important keywords and sentences. The input is the collected text data, and the output is text data in a clean and consistent format.
[0260] Step 3:
[0261] The server trains a generative AI model using preprocessed data. Specifically, it uses a model such as GPT-3, training the model by splitting the data into training and test sets. It optimizes the model parameters using backpropagation. The input is preprocessed text data, and the output is the trained AI model.
[0262] Step 4:
[0263] The server generates a virtual leader emulator using a trained AI model. This emulator has speech recognition and speech synthesis capabilities and is deployed on the cloud. The input is the trained AI model, and the output is the virtual leader emulator.
[0264] Step 5:
[0265] The user navigates a virtual environment through a virtual reality device (HMD) and selects products. They use an interface to interact with a virtual instructor emulator and ask questions about the products. Input is either voice or text input from the user, and output is a voice response from the virtual instructor emulator.
[0266] Step 6:
[0267] The server collects interaction data between the user and the virtual instructor emulator. Specifically, it collects data including dialogue history and usage patterns, and uses this data to retrain the AI model. The input is interaction data, and the output is the retrained AI model.
[0268] Step 7:
[0269] The server deploys a retrained AI model to the cloud, improving the accuracy and responsiveness of the virtual mentor emulator. This allows users to receive up-to-date information and optimal advice in real time. The input is the retrained AI model, and the output is the improved virtual mentor emulator.
[0270] By following these steps, a system will be completed that improves the efficiency of business operations and the accuracy of decision-making.
[0271] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0272] System Overview
[0273] This invention relates to a system that generates a virtual CEO emulator by collecting and preprocessing information related to the CEO, and training a generative AI model and emotion engine, enabling interaction through a user interface. This system can recognize the user's emotions and adjust its responses based on those emotions.
[0274] Program Processing Overview
[0275] This system consists of the following main phases:
[0276] 1. Data Collection Phase
[0277] 2. Data preprocessing phase
[0278] 3. Model learning phase
[0279] 4. Sentiment engine learning phase
[0280] 5. CEO emulator generation phase
[0281] 6. Interaction phase
[0282] Operation description of the program
[0283] Data collection phase
[0284] The server collects data related to the CEO from the specified websites or SNS platforms using web scraping technology. It targets blog articles, interviews, SNS posts, etc. of the CEO published on the Internet.
[0285] The user uploads additional data related to the CEO, such as internal materials and past email data, using the system interface.
[0286] Data storage
[0287] The server stores the collected data in a database for unified management. The data to be stored includes formats such as text, images, and audio files.
[0288] Noise removal and data cleansing
[0289] The server performs noise removal processing on the collected data. It executes operations such as correcting spelling mistakes and deleting unnecessary tags and special symbols.
[0290] The server unifies the data format and prepares the text data in a clean state. For example, it deletes HTML tags and converts them to plain text.
[0291] Data analysis using natural language processing
[0292] The server uses natural language processing techniques to analyze clean text data. It performs tokenization, part-of-speech tagging, and sentence analysis to extract important keywords and sentences.
[0293] Creating a dataset
[0294] The server splits the analyzed data into a training set and a test set. For example, 70% of the data might be used for the training set and 30% for the test set.
[0295] Training of generative AI models
[0296] The server trains a generative AI model using a training set. It uses deep learning algorithms and performs backpropagation to optimize the model's parameters.
[0297] The server trains the model multiple times, specifying the number of epochs, to improve the model's accuracy.
[0298] Learning the Emotion Engine
[0299] The server trains the emotion engine using a training set. It analyzes user voice and text input data, assigns emotion labels, and trains the model.
[0300] The server will implement a function that uses an emotion engine to recognize emotions in real time from user input.
[0301] Model evaluation
[0302] The server evaluates the performance of the trained generative AI models and sentiment engines using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[0303] The server performs retraining as needed to further improve the performance of the model.
[0304] Generation of the CEO Emulator
[0305] The server uses the final generative AI model and the emotion engine to generate a virtual CEO emulator for a two-dimensional interface. This emulator has speech recognition and speech synthesis functions and can interact with users in real time.
[0306] Deployment of the User Interface
[0307] The server deploys the generated virtual CEO emulator on the cloud. It provides a URL that can be accessed by users through a browser and functions as a user interface.
[0308] Initiation of Interaction with the User
[0309] The user uses the web interface to initiate interaction with the virtual CEO. Questions and consultations are made using text input or voice input.
[0310] The server analyzes the user's input in real time and generates an appropriate response. The generated response is provided to the user as voice output using a speech synthesis engine.
[0311] The server analyzes the user's emotions with the emotion engine and adjusts the response based on the recognized emotions. For example, if the user is angry, the response is given in a gentle tone.
[0312] Collection and Retraining of Interaction Data
[0313] The server collects the interaction logs with the user and saves them as interaction data. This enables continuous improvement and accuracy enhancement of the system.
[0314] The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue.
[0315] Specific example
[0316] Data collection examples
[0317] The server collects the CEO's social media posts and interview articles from the past five years and stores them in a database.
[0318] Users can supplement the collected data by uploading personal emails and internal documents to the system.
[0319] Data preprocessing example
[0320] The server analyzes the collected data and corrects typos and grammatical errors. It also removes unnecessary tags and special characters, converting the data into clean text.
[0321] Model Learning Example
[0322] The server uses pre-processed data to train generative AI models and emotion engines, improving their ability to understand the CEO's speaking patterns and users' emotions.
[0323] Emulator generation example
[0324] The server generates a 2D virtual CEO and implements speech recognition and synthesis capabilities. These emulators are deployed on the cloud as the user interface.
[0325] User interaction examples
[0326] Users ask questions about project progress to a virtual CEO and receive specific advice. The server analyzes the user's questions and emotions in real time to generate the most appropriate answers.
[0327] Based on the above, the present invention provides a system that allows CEOs to leverage their knowledge and judgment into the future. This system recognizes user emotions and adjusts responses based on those emotions, enabling more natural and effective dialogue. This, in turn, improves the efficiency of corporate operations and the accuracy of decision-making.
[0328] The following describes the processing flow.
[0329] Step 1: Data Collection
[0330] The server collects data related to the CEO from specified websites and social media platforms using web scraping techniques. Specifically, it uses APIs to obtain social media posts and crawling techniques to obtain text data from blog posts and interviews.
[0331] Users upload additional CEO-related data, such as internal company documents and past email data, using the system interface. For example, they can upload files using a drag-and-drop function.
[0332] Step 2: Save
[0333] The server stores the collected data in a database for centralized management. For example, it might use a NoSQL database to store data in different formats (text, images, audio files, etc.).
[0334] Step 3: Noise Reduction and Data Cleansing
[0335] The server performs noise reduction on the collected data. This is done using automated scripts and natural language processing tools to correct spelling errors, remove unnecessary tags and special characters, and so on.
[0336] The server standardizes data formats and cleans up text data. For example, it removes HTML tags and converts it to plain text.
[0337] Step 4: Data analysis using natural language processing
[0338] The server uses natural language processing techniques to analyze clean text data. Specifically, it performs processes such as tokenization, part-of-speech tagging, sentence analysis, and semantic analysis.
[0339] The server extracts important keywords and sentences from the analysis results and stores them in a database. For example, it may use text frequency analysis or TF-IDF scoring.
[0340] Step 5: Create a dataset
[0341] The server splits the analyzed data into a training set and a test set. For example, 70% of the data could be used for the training set and 30% for the test set. The script is then executed to distribute the data randomly.
[0342] Step 6: Training the Generative AI Model
[0343] The server trains a generative AI model (e.g., GPT-3 or BERT) using the training set. It then performs backpropagation to optimize the model's parameters using a deep learning algorithm.
[0344] The server trains the model multiple times, specifying the number of epochs (the number of training iterations), to improve the model's accuracy. Batch processing is used to perform calculations efficiently during this process.
[0345] Step 7: Learning the Emotion Engine
[0346] The server simultaneously trains the emotion engine using the training set. It analyzes user voice and text input data, assigns emotion labels, and trains the model. Specifically, it performs supervised learning using labeled datasets.
[0347] The server implements a function to recognize emotions in real time from user input using an emotion engine. For example, it uses speech feature extraction for emotion analysis from speech and emotion classification algorithms for emotion analysis from text.
[0348] Step 8: Model Evaluation
[0349] The server evaluates the performance of the trained generative AI models and sentiment engines using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[0350] The server retrains the model as needed to further improve its performance, for example, by performing hyperparameter tuning.
[0351] Step 9: Generate the CEO emulator
[0352] The server uses the final generative AI model and emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time. A graphical user interface (GUI) is designed to be user-friendly.
[0353] Step 10: Deploying the User Interface
[0354] The server deploys the generated virtual CEO emulator to the cloud. For example, it uses a web application framework to provide a URL that users can access through a browser, thus functioning as a user interface.
[0355] Step 11: Start interacting with the user
[0356] Users initiate a conversation with a virtual CEO using a web interface. They can ask questions and seek advice using text or voice input. The service operates in a browser and uses a microphone for voice input.
[0357] The server analyzes user input (text or voice) in real time and generates an appropriate response. The generated response is then delivered to the user as voice output using a speech synthesis engine. For example, natural language generation (NLG) technology is used to generate text, and a text-to-speech (TTS) engine is used to generate speech.
[0358] The server analyzes the user's emotions using an emotion engine and adjusts its response based on the recognized emotions. For example, if the user is angry, it will respond in a calm tone; if the user is sad, it will respond in a comforting tone.
[0359] Step 12: Collect interaction data and retrain
[0360] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement. The interaction logs stored in the database include user questions, virtual CEO responses, and user sentiment information.
[0361] The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue. Regular retraining allows the system to evolve over time, enabling more accurate responses.
[0362] (Example 2)
[0363] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0364] In current corporate operations, there is a lack of means to virtually replicate the knowledge and judgment of the CEO and to facilitate effective communication with employees and stakeholders. Furthermore, there is no system that can recognize user emotions in real time and adjust responses accordingly. This can potentially lead to a decrease in the efficiency of decision-making and problem-solving within companies.
[0365] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0366] In this invention, the server includes means for collecting information related to the CEO, means for preprocessing the collected information to remove noise and standardize the data format, means for analyzing the preprocessed data using natural language processing techniques to extract important keywords and sentences, means for training a generative AI model using the extracted data, means for generating a virtual CEO emulator with emotion recognition capabilities using the generated AI model, means for interacting with the virtual CEO through a user interface, means for collecting interaction data and continuously retraining the AI model, means for recognizing emotions in real time from user input, and means for adjusting responses based on recognized emotions. This makes it possible to virtually reproduce the knowledge and judgment of the CEO and generate appropriate responses according to the user's emotions.
[0367] "CEO-related information" refers to text, image, and audio data related to the CEO, such as blog posts, interviews, social media posts, internal documents, and past email data.
[0368] "Means of collection" refers to systems that include methods for obtaining specified information from the internet using web scraping technology, as well as methods for uploading internal documents and past email data via a user interface.
[0369] "Methods for preprocessing, noise reduction, and data format standardization" refer to techniques for processing collected data to create a clean data format by correcting spelling errors, removing unnecessary tags and special characters, and converting HTML tags to plain text.
[0370] "Natural language processing technology" refers to algorithms and methods used to analyze text data and extract important keywords and sentences through tokenization, part-of-speech tagging, and sentence analysis.
[0371] A "generative AI model" is an AI model trained using deep learning algorithms that has the ability to generate appropriate outputs for specific inputs.
[0372] A "virtual CEO emulator" is a virtual character created based on a generative AI model, possessing speech recognition and speech synthesis capabilities, and serving as an interface that can interact with users in real time.
[0373] A "user interface" refers to the graphical screen or web page that a user uses to interact with a virtual CEO.
[0374] "Interaction data" refers to digital information, including dialogue logs and response data from interactions between the user and the virtual CEO emulator.
[0375] "Emotion recognition functionality" refers to technology that analyzes user voice and text input data and identifies the user's emotions based on emotion labels extracted from that data.
[0376] "Means of adjusting responses based on emotions" refers to methods or algorithms for adjusting the content and tone of generated responses according to the user's emotions identified by emotion recognition functions.
[0377] This invention is a system that collects information related to a CEO, trains a generative AI model based on that information to generate a virtual CEO emulator, and allows the user to interact with it through an interface. This system also has the ability to recognize the user's emotions and adjust its responses based on those emotions. The specific embodiments of this system are described in detail below.
[0378] Data collection and storage
[0379] The server collects information related to the CEO from specified websites and social media platforms. This task utilizes web scraping techniques such as BeautifulSoup and Scrapy. The collected information includes the CEO's blog posts, interviews, and social media posts. The collected data is stored in databases such as MySQL® and MongoDB.
[0380] Specific example:
[0381] The server collects social media posts and interview articles from the past five years and stores them in a database.
[0382] Users upload internal documents and past email data using the system interface.
[0383] Data preprocessing
[0384] The server performs noise reduction and data formatting on the collected data. Specifically, it corrects spelling errors using regular expressions and removes unnecessary HTML tags and special characters. Text data is converted to plain text. Python is often used for this process.
[0385] Analysis using natural language processing
[0386] The server applies natural language processing techniques to the pre-processed data. For example, NLTK or SpaCy are used to tokenize text data, tag parts of speech, and analyze sentences, extracting important keywords and sentences.
[0387] Specific example:
[0388] The server tokenizes the text data, tags it with parts of speech, and extracts important keywords.
[0389] Model training
[0390] The server trains generative AI models using the extracted data. The deep learning frameworks used include TENSORFLOW® and PyTorch. Parameter optimization is performed using backpropagation with the training set.
[0391] Learning the Emotion Engine
[0392] The server trains an emotion engine, as well as a trained generative AI model. This emotion engine uses a BERT-based emotion analysis model to extract emotion labels (e.g., joy, sadness, anger) from the user's voice and text input data.
[0393] Specific example:
[0394] The server analyzes the audio data and trains a model to assign emotion labels.
[0395] Creating a virtual emulator
[0396] The server integrates a generative AI model and an emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition (Google Speech-to-Text API) and speech synthesis (Amazon Polly) capabilities.
[0397] Deployment and User Interaction
[0398] The server deploys the generated virtual CEO emulator to a cloud environment (such as AWS or Azure®) and makes it accessible to users via a browser. Users can initiate an interaction with the virtual CEO using the interface. The server analyzes questions entered via text or voice in real time and generates appropriate responses. These responses are output as speech using a speech synthesis engine. It also features an emotion engine that analyzes the user's emotions and adjusts the tone of the responses accordingly.
[0399] Specific example:
[0400] The user enters, "Based on the CEO's latest social media posts, please share your insights into current market trends."
[0401] The server analyzes this input and generates a response through a virtual CEO emulator.
[0402] Collection of interaction data and retraining
[0403] The server collects user interaction logs and stores them as interaction data. Based on this data, the model is retrained to continuously improve performance.
[0404] The above describes a specific embodiment for implementing the present invention, which will lead to increased efficiency in decision-making and problem-solving within companies.
[0405] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0406] Step 1: Data Collection
[0407] The server collects information related to the CEO from specified websites and social media platforms. Specifically, it uses web scraping techniques (e.g., BeautifulSoup, Scrapy) to obtain the data.
[0408] Users upload internal documents and past email data through the system interface.
[0409] Input: Website URL, social media platform account information, user-uploaded internal documents and email data.
[0410] Output: Collected text data, image data, and audio data.
[0411] Step 2: Save Data
[0412] The server stores the collected data in a database (e.g., MySQL, MongoDB). Text data, image data, audio data, etc., are stored in the appropriate format according to their respective requirements.
[0413] Input: Collected text data, image data, and audio data.
[0414] Output: Clean data stored in the database.
[0415] Step 3: Noise Reduction and Data Cleansing
[0416] The server performs noise reduction processing on the collected data. Specifically, it uses regular expressions to correct spelling mistakes and remove unnecessary HTML tags and special characters.
[0417] Input: Raw data stored in the database.
[0418] Output: Denoised and cleaned text data.
[0419] Step 4: Data analysis using natural language processing
[0420] The server uses pre-processed data to apply natural language processing techniques (e.g., NLTK, SpaCy) to extract important keywords and sentences through tokenization, part-of-speech tagging, and sentence analysis.
[0421] Input: Clean text data.
[0422] Output: Tokenized keywords, part-of-speech tags, and parsed sentences.
[0423] Step 5: Create a dataset
[0424] The server splits the analyzed data into a training set (e.g., 70% of the data) and a test set (e.g., 30% of the data). These datasets are used for model training and evaluation.
[0425] Input: Analyzed keywords and sentences.
[0426] Output: Training set, test set.
[0427] Step 6: Training the Generative AI Model
[0428] The server trains a generative AI model (e.g., GPT-3) using a training set. Deep learning frameworks used include TensorFlow and PyTorch. Parameter optimization is performed using backpropagation.
[0429] Input: Training set.
[0430] Output: A trained generative AI model.
[0431] Step 7: Learning the Emotion Engine
[0432] The server trains an emotion engine (e.g., a BERT-based emotion analysis model) using a training set. The model is trained by assigning emotion labels to user voice and text input data.
[0433] Input: Training set, user voice and text data.
[0434] Output: Trained emotion engine.
[0435] Step 8: Model Evaluation
[0436] The server evaluates the performance of generative AI models and emotion engines trained using a test set. Evaluation metrics include accuracy, recall, and F1 score, and retraining is performed as needed.
[0437] Input: Test set.
[0438] Output: Evaluation results (accuracy, recall, F1 score).
[0439] Step 9: Generate the CEO emulator
[0440] The server integrates the final generative AI model and emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition (Google Speech-to-Text API) and speech synthesis (Amazon Polly) capabilities.
[0441] Input: Trained generative AI model, emotion engine.
[0442] Output: Virtual CEO emulator.
[0443] Step 10: Deploying the User Interface
[0444] The server deploys the generated virtual CEO emulator to a cloud environment (AWS or Azure) and makes it accessible to users through a browser.
[0445] Input: Virtual CEO emulator.
[0446] Output: Deployed interface (URL).
[0447] Step 11: Interacting with the user
[0448] Users initiate a conversation with a virtual CEO through their browser, entering questions and concerns via text or voice.
[0449] The server analyzes user input in real time and generates an appropriate response. The generated response is delivered to the user using a speech synthesis engine. Additionally, an emotion engine analyzes the user's emotions and adjusts the tone of the response accordingly.
[0450] Input: User text and voice input.
[0451] Output: Virtual CEO's response (text and audio).
[0452] Step 12: Collect interaction data and retrain
[0453] The server collects user interaction logs and stores them in a database as interaction data. The collected data is used to retrain the emotion engine and generative AI models.
[0454] Input: Dialogue log between the user and the virtual CEO.
[0455] Output: Collected dialogue logs, retrained AI model, and emotion engine.
[0456] (Application Example 2)
[0457] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0458] In modern factory operations, diverse data needs to be collected and analyzed in real time, but systems for handling this data efficiently and effectively are not yet widespread. Furthermore, there is a lack of means to provide appropriate instructions and advice in real time based on worker emotions and production status, which can lead to problems such as decreased production efficiency and increased worker stress.
[0459] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting information related to the CEO, means for preprocessing the collected information to remove noise and unify the data format, means for training a generative AI model using the preprocessed data to generate a virtual CEO emulator, means for interacting with the virtual CEO through a user interface, means for collecting interaction data and continuously retraining the AI model, means for collecting and monitoring data on each machine and worker in the factory in real time, and means for analyzing production status and worker emotions based on the collected data, and for the virtual CEO to provide appropriate instructions and advice. This makes it possible to improve the efficiency of factory operations and reduce worker stress.
[0460] "Information related to the CEO" refers to publicly available information, internal documents, past email data, blog posts, interviews, and social media posts concerning the company's chief executive officer.
[0461] "Preprocessing" refers to a series of processes to remove noise from collected data and standardize the data format.
[0462] "Noise reduction" refers to the process of removing unnecessary elements (such as spelling mistakes, special characters, and HTML tags) from collected data.
[0463] "Data format standardization" refers to the process of converting data in different formats into a consistent format.
[0464] A "generative AI model" refers to an artificial intelligence model that has been trained using pre-collected data and is capable of generating appropriate responses to specific tasks.
[0465] A "virtual CEO emulator" refers to a virtual character based on a generative AI model trained to mimic the speech patterns and decision-making styles of a CEO.
[0466] "User interface" refers to the interactive means by which a user interacts with a system, including browsers, displays, etc.
[0467] "Interaction data" refers to data that records the dialogue logs and responses between the user and the virtual CEO emulator.
[0468] "Continuously retraining an AI model" refers to the process of periodically retraining an AI model to improve its accuracy based on collected interaction data.
[0469] "Data on each machine and worker within the factory" refers to real-time data such as the operating status of factory production equipment, the behavior of workers, and their emotional states.
[0470] "Real-time data collection and monitoring" refers to the process of instantly acquiring current conditions and continuously monitoring that data.
[0471] "Production status" refers to information related to the progress of product production within the factory and the operating status of machinery.
[0472] "Worker emotions" refers to data related to emotions, such as workers' stress levels, satisfaction levels, and fatigue levels.
[0473] "Providing instructions and advice" refers to the process by which the virtual CEO emulator recommends specific actions and measures based on the data it has collected.
[0474] This invention relates to a system using a virtual CEO emulator aimed at improving the efficiency of factory operations and reducing worker stress. The system has the following configuration:
[0475] Hardware and software to be used
[0476] The server will use a high-performance server, natural language processing libraries (such as SpaCy and NLTK), deep learning frameworks (such as TensorFlow and PyTorch), sentiment recognition libraries (such as openSMILE), and web scraping tools (such as BeautifulSoup and Scrapy).
[0477] The devices include IoT sensors within the factory, wearable devices for workers, tablets, and displays.
[0478] Program processing flow
[0479] Data collection
[0480] The server collects data in real time from IoT sensors and wearable devices within the factory. It also periodically collects relevant industry news and technical literature using web scraping tools.
[0481] Specific example:
[0482] "Workers' smartwatches measure heart rate and stress levels and transmit the data to a cloud server in real time."
[0483] "IoT sensors monitor the operating status of each machine and transmit the data to a server."
[0484] Data preprocessing
[0485] The server removes noise from the collected data and standardizes the data format. Specifically, it corrects typos and removes unnecessary tags and special characters.
[0486] Model Learning
[0487] The server uses pre-processed data to train generative AI models and emotion engines. This allows a virtual CEO emulator to understand the emotional state and productivity of workers and provide appropriate instructions and advice.
[0488] Generating a virtual CEO emulator
[0489] The server generates a virtual CEO emulator using the generated generative AI model and emotion engine. This emulator can interact in real time with workers and managers via tablets or displays.
[0490] Example of a prompt:
[0491] "Could you tell me the current status of the production line?"
[0492] "Monitor the stress levels of the workers and suggest breaks if necessary."
[0493] User Interface and Interaction
[0494] Users can receive real-time feedback and advice on production status and individual issues through interaction with a virtual CEO emulator. The server has the capability to analyze user input and emotions and generate appropriate responses.
[0495] Specific example:
[0496] "Workers ask a virtual CEO questions about the project's progress and receive specific advice. The server analyzes the user's questions and emotions in real time to generate the most appropriate response."
[0497] Collection of interaction data and retraining
[0498] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement. The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue.
[0499] In this way, the present invention improves the efficiency of factory operations and reduces stress on workers.
[0500] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0501] Step 1:
[0502] The server collects data in real time from IoT sensors and wearable devices within the factory. Inputs include data such as the operating status of each machine, workers' heart rates, and stress levels. Based on this, the server centrally stores this data in a database, making it immediately accessible.
[0503] Step 2:
[0504] The server preprocesses the collected data. The input is the raw data collected in step 1. Specifically, it corrects typos and grammatical errors, removes HTML tags and unnecessary special characters, and standardizes the data into a clean text format. The output is the clean data saved in the standardized format.
[0505] Step 3:
[0506] The server applies natural language processing techniques to the pre-processed data to extract important keywords and sentences. The input is the clean text data from step 2. Specific operations include tokenization, part-of-speech tagging, and sentence analysis. The output is a list of the extracted keywords and sentences.
[0507] Step 4:
[0508] The server trains a generative AI model using preprocessed data and natural language processing results. The input is the data from steps 2 and 3. A deep learning framework (such as TensorFlow or PyTorch) is used, and backpropagation is performed to optimize the model parameters. The output is the trained generative AI model.
[0509] Step 5:
[0510] The server trains its emotion engine using an emotion recognition library (such as openSMILE). The input consists of worker voice and text data. Its specific actions include voice analysis and emotion labeling. The output is the trained emotion engine.
[0511] Step 6:
[0512] The server generates a virtual CEO emulator using a trained generative AI model and an emotion engine. The input is the trained model from steps 4 and 5. The virtual CEO emulator can interact with the user in real time via a tablet or display. The output is the virtual CEO emulator running on the user interface.
[0513] Step 7:
[0514] The user begins interacting with a virtual CEO emulator. Specifically, they input prompts such as, "Please tell me the current status of the production line," or "Please monitor the stress levels of the workers and suggest breaks if necessary." The input can be text or voice from the user. The server analyzes this and generates the most appropriate response. The output is appropriate advice or instructions provided to the user.
[0515] Step 8:
[0516] The server collects user interaction data and stores it in the interaction database. The input is the interaction log from step 7. The output is the stored interaction data.
[0517] Step 9:
[0518] The server retrains the generative AI model and emotion engine using the collected interaction data. The input is the interaction data from step 8. As a result of the retraining, the accuracy of the AI model and emotion engine improves, enabling more natural and effective dialogue. The output is the further improved generative AI model and emotion engine.
[0519] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0520] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0521] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0522] [Second Embodiment]
[0523] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0524] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0525] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0526] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0527] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0528] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0529] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0530] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0531] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0532] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0533] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0534] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0535] System Overview
[0536] This invention relates to a system that generates a virtual CEO emulator by collecting and preprocessing information related to the CEO and training a generative AI model, enabling interaction through a user interface. This system can learn the CEO's statements and decisions, imitate them in real time, and interact with the user via a two-dimensional interface.
[0537] Program Processing Overview
[0538] This system consists of the following main phases:
[0539] 1. Data Collection Phase
[0540] 2. Data preprocessing phase
[0541] 3. Model training phase
[0542] 4. CEO Emulator Generation Phase
[0543] 5. Interaction Phase
[0544] Program operation description
[0545] Data collection phase
[0546] The server collects publicly available information related to the CEO using web scraping techniques. This includes social media posts, blog articles, interviews, and publicly available emails.
[0547] Users can improve the accuracy of data collection by uploading internal documents and files to the system.
[0548] Data preprocessing phase
[0549] The server preprocesses the collected data, removing noise and standardizing the data format. Specifically, it uses a grammar checking tool to correct typos and removes unnecessary HTML tags and special characters.
[0550] The server uses natural language processing technology to analyze text data and extract important keywords and sentences.
[0551] Model Learning Phase
[0552] The server uses a generative AI model (e.g., GPT-3 or BERT) to train on preprocessed data. Initial setup is performed, and the model is instantiated.
[0553] The server splits the data into a training set and a test set and performs training. During this process, it optimizes the model parameters using backpropagation.
[0554] The server evaluates the model's performance using a test set and retrains it as needed.
[0555] CEO Emulator Generation Phase
[0556] The server generates a virtual CEO emulator as a 2D interface. This emulator is equipped with speech recognition and speech synthesis capabilities, enabling real-time interaction with the user.
[0557] The server deploys the generated interface as a user interface on the cloud, making it accessible to users.
[0558] Interaction Phase
[0559] Users interact with a virtual CEO through their device. For example, they log in to a web application using a browser and open a dialogue screen.
[0560] The server analyzes user input (text or voice) and generates an appropriate response in real time. The generated response is then delivered to the user as voice output through a speech synthesis engine.
[0561] The server continuously collects new dialogue data and retrains the model, thereby improving the system's accuracy and response quality.
[0562] Specific example
[0563] Data collection examples
[0564] The server collects the CEO's social media posts and interview articles from the past five years and stores them in a database.
[0565] Users can supplement the collected data by uploading personal emails and internal documents to the system.
[0566] Data preprocessing example
[0567] The server analyzes the collected data and corrects typos and grammatical errors. It also removes unnecessary tags and special characters, converting the data into clean text.
[0568] Model Learning Example
[0569] The server uses pre-processed data to train a generative AI model, learning the CEO's speaking patterns. For example, it incorporates frequently used phrases and decision-making tendencies into the model.
[0570] Emulator generation example
[0571] The server generates a 2D virtual CEO and implements speech recognition and synthesis capabilities. These emulators are deployed on the cloud as the user interface.
[0572] User interaction examples
[0573] Users ask questions about project progress to a virtual CEO and receive specific advice. The server analyzes the user's questions in real time and generates the most appropriate answers.
[0574] In summary, this invention provides a system that allows CEOs to leverage their knowledge and judgment for the future. This will lead to increased efficiency in corporate operations and improved accuracy in decision-making.
[0575] The following describes the processing flow.
[0576] Step 1: Data Collection
[0577] The server uses web scraping techniques to collect data related to the CEO from specified websites and social media platforms. This includes publicly available blog posts, interviews, and social media posts by the CEO.
[0578] Users can use the system interface to upload additional CEO-related data, such as internal documents and past email data.
[0579] Step 2: Save
[0580] The server stores the collected data in a database for centralized management. The stored data includes various formats such as text, images, and audio files.
[0581] Step 3: Noise Reduction and Data Cleansing
[0582] The server performs noise reduction processing on the collected data. This includes tasks such as correcting spelling errors and removing unnecessary tags and special characters.
[0583] The server standardizes data formats and cleans up text data. For example, it removes HTML tags and converts it to plain text.
[0584] Step 4: Data analysis using natural language processing
[0585] The server uses natural language processing techniques to analyze clean text data. It performs tokenization, part-of-speech tagging, and sentence analysis to extract important keywords and sentences.
[0586] Step 5: Create a dataset
[0587] The server splits the analyzed data into a training set and a test set. For example, 70% of the data might be used for the training set and 30% for the test set.
[0588] Step 6: Training the Generative AI Model
[0589] The server trains a generative AI model using a training set. It uses deep learning algorithms and performs backpropagation to optimize the model's parameters.
[0590] The server trains the model multiple times, specifying the number of epochs, to improve the model's accuracy.
[0591] Step 7: Model Evaluation
[0592] The server evaluates the performance of the trained model using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[0593] The server retrains the model as needed to further improve its performance.
[0594] Step 8: Generating the CEO Emulator
[0595] The server uses the final generative AI model to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time.
[0596] Step 9: Deploying the User Interface
[0597] The server deploys the generated virtual CEO emulator to the cloud. It provides a URL that users can access via a browser, and functions as the user interface.
[0598] Step 10: Start interacting with the user
[0599] Users initiate a conversation with a virtual CEO using a web interface. They can ask questions and seek advice using text or voice input.
[0600] The server analyzes user input in real time and generates an appropriate response. The generated response is then delivered to the user as speech output using a speech synthesis engine.
[0601] Step 11: Collect interaction data and retrain
[0602] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement.
[0603] The server uses the collected interaction data to retrain the model and improve the quality of the dialogue.
[0604] (Example 1)
[0605] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0606] In today's business environment, there is a growing need to quickly and accurately mimic the judgments and instructions of managers in order to improve the efficiency of corporate operations and the precision of decision-making. Furthermore, there is an increasing need for systems that can utilize the knowledge and experience of managers even when they are physically absent. However, conventional technologies do not adequately provide methods for mimicking managers' thought patterns and statements in real time.
[0607] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0608] In this invention, the server includes means for collecting publicly available information, means for preprocessing the collected information, performing grammar checks and unifying data formats, means for training a generative AI model using the preprocessed data to generate a virtual manager emulator, means for interacting with the virtual manager through a user interface, and means for collecting interaction data and continuously retraining the AI model. This enables real-time imitation of the manager's knowledge and judgment, improving the efficiency of business operations and the accuracy of decision-making.
[0609] "Public information" refers to information that is accessible to a large number of people, such as information posted on websites, social media posts, blog articles, interview articles, and publicly available emails.
[0610] "Preprocessing" is the process of correcting typographical errors in collected data, removing unnecessary HTML tags and special characters, and standardizing data formats.
[0611] A "generative AI model" is an artificial intelligence model that is trained using collected and pre-processed data, and includes, for example, natural language generation models such as GPT-3.
[0612] A "virtual executive emulator" is a virtual character that mimics the statements and decisions of executives in real time based on a generated AI model, and it has speech recognition and speech synthesis capabilities.
[0613] "User interface" refers to the screens and applications that allow a user to interact with a virtual manager emulator, and includes web browsers and dedicated applications.
[0614] "Interaction data" refers to the data from conversations between the user and a virtual manager emulator, and this data is used to retrain the AI model.
[0615] "Natural language processing technology" refers to techniques for analyzing text data and extracting important keywords and sentences, and includes libraries such as NLTK and SpaCy.
[0616] "Speech recognition" is a technology that converts speech data into text data, and Google Speech-to-Text is an example of this.
[0617] "Speech synthesis" is a technology that converts text data into speech data, and Amazon Polly is an example of this.
[0618] This invention is a system that generates a virtual CEO emulator by collecting and pre-processing information related to the CEO and training a generated AI model, enabling interaction through a user interface. This system imitates the judgments and statements of executives in real time based on publicly available information, thereby improving the efficiency of corporate operations and the accuracy of decision-making.
[0619] The system configuration is mainly described in the following phases:
[0620] 1. Data Collection Phase
[0621] 2. Data preprocessing phase
[0622] 3. Model training phase
[0623] 4. CEO Emulator Generation Phase
[0624] 5. Interaction Phase
[0625] Data collection phase
[0626] The server uses web scraping techniques (e.g., Beautiful Soup) to collect publicly available information such as social media posts, blog articles, interviews, and published emails. Specifically, it uses Python libraries to crawl websites containing certain keywords (e.g., names of business owners, company names).
[0627] Users can improve the quality and quantity of data collection by uploading internal documents and personal emails to the system. The uploaded files are stored in a database on the server.
[0628] Data preprocessing phase
[0629] The server corrects the collected data using a grammar checking tool (e.g., Grammarly API) to eliminate typos and grammatical errors.
[0630] The server removes unnecessary HTML tags and special characters from the collected data, converting it into clean text data.
[0631] The server analyzes text data using natural language processing technologies (e.g., NLTK, SpaCy) to extract important keywords and sentences. For example, it can extract key topics and keywords from interview articles with business executives.
[0632] Model Learning Phase
[0633] The server instantiates a generative AI model (e.g., GPT-3) and loads pre-processed data. Instances of the generative AI model are created using the OpenAI API.
[0634] The server splits the preprocessed data into a training set and a test set, and then trains the model.
[0635] The server uses backpropagation to optimize the parameters of the generated AI model and retrains it to improve its accuracy.
[0636] CEO Emulator Generation Phase
[0637] The server generates a virtual manager emulator using a pre-trained generative AI model. This emulator is designed as a two-dimensional interface and includes speech recognition (e.g., Google Speech-to-Text) and speech synthesis (e.g., Amazon Polly) capabilities.
[0638] The server deploys a virtual manager emulator on the cloud and makes it accessible to users. For example, it might be deployed on AWS and configured to be accessible to users via a browser.
[0639] Interaction Phase
[0640] The user logs into the web application from their device and opens the interactive interface.
[0641] Example: A user logs into the system using a browser and asks a virtual manager, "What is the current project status?"
[0642] The server analyzes user input (text or voice) and generates an appropriate response.
[0643] Example: The server analyzes the user's question and provides the generated response as speech output using a speech synthesis engine.
[0644] The server continuously collects new interaction data and retrains the generative AI model. This improves the system's accuracy and response quality.
[0645] Examples of specific prompt statements include the following:
[0646] "Could you tell me the CEO's view on the progress of the current project?"
[0647] "We'd like to hear the CEO's opinion on our new marketing strategy."
[0648] As described above, the present invention makes it possible to mimic the knowledge and judgment of managers in real time, thereby improving the efficiency of corporate operations and the accuracy of decision-making.
[0649] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0650] The specific processing flow of the program
[0651] Step 1: Data Collection
[0652] The server launches a web scraping script to collect publicly available information such as social media posts, blog articles, interviews, and published emails. Specifically, it uses the Python Beautiful Soup library to crawl websites containing specific keywords.
[0653] Input: Website URL or specific keywords
[0654] Output: Collected text data (e.g., text from social media posts or blog articles)
[0655] Specific operation: The server accesses a specific website, retrieves the page content, and extracts the necessary parts.
[0656] Step 2: Data Upload
[0657] Users upload internal documents and personal emails to the system. The uploaded files are stored in a database on the server.
[0658] Input: A file selected by the user (e.g., PDF, text file)
[0659] Output: Text data stored in the database
[0660] Specific operation: The user sends a file to the server using the system's upload form, and the server saves its contents to the database.
[0661] Step 3: Data Cleaning
[0662] The server uses a grammar checking tool (e.g., Grammarly API) to correct typos and grammatical errors in the collected data.
[0663] Input: Collected text data
[0664] Output: Clean text data with typos and grammatical errors corrected.
[0665] Specific operation: The server sends text data to a grammar checking tool and receives the corrected text.
[0666] Step 4: Unify text formatting
[0667] The server removes unnecessary HTML tags and special characters from the collected data to standardize the data format.
[0668] Input: Text data with typos and grammatical errors corrected.
[0669] Output: Clean text data
[0670] Specific operation: The server uses regular expressions to remove HTML tags and special characters, converting the data into clean text.
[0671] Step 5: Natural Language Processing (NLP)
[0672] The server uses natural language processing techniques (e.g., NLTK, SpaCy) to analyze text data and extract important keywords and sentences.
[0673] Input: Clean text data
[0674] Output: Extracted keywords and important sentences
[0675] Specific operation: The server starts a natural language processing library, analyzes the text data, and extracts important information.
[0676] Step 6: Model instantiation and data loading
[0677] The server instantiates a generative AI model (e.g., GPT-3) and loads pre-processed data.
[0678] Input: Extracted keywords and important sentences
[0679] Output: Text data loaded as training data
[0680] Specific operation: The server uses the OpenAI API to create a GPT-3 instance and load the pre-processed data.
[0681] Step 7: Model Training
[0682] The server divides the pre-processed data into a training set and a test set, and then trains the generative AI model.
[0683] Input: Loaded training data
[0684] Output: Trained generative AI model
[0685] Specific operation: The server optimizes the model using training data and evaluates its accuracy using test data.
[0686] Step 8: Generate CEO emulator
[0687] The server generates a virtual manager emulator using a pre-trained generative AI model. This emulator has a two-dimensional interface and features speech recognition and speech synthesis capabilities.
[0688] Input: Trained generative AI model
[0689] Output: Virtual manager emulator
[0690] Specific operation: The server generates characters from the model and integrates speech recognition and speech synthesis engines.
[0691] Step 9: Interface Deployment
[0692] The server deploys a virtual manager emulator on the cloud, making it accessible to users.
[0693] Input: Virtual manager emulator
[0694] Output: Deployed user interface
[0695] Specific operation: The server deploys the emulator to a cloud service such as AWS and configures it so that users can access it via a browser.
[0696] Step 10: User Interaction
[0697] The user logs into the web application from their device, opens a conversational interface, and interacts with a virtual manager.
[0698] Input: User voice or text input
[0699] Output: Response from the virtual manager
[0700] Specific operation: The user inputs a question or instruction, and the server generates a response and sends it back.
[0701] Step 11: Data Collection and Retraining
[0702] The server continuously collects conversational data between the user and the virtual manager, and uses it to retrain the generated AI model.
[0703] Input: User interaction data
[0704] Output: Retrained generative AI model
[0705] Specific operation: The server collects interaction data, periodically retrains the model, and improves the system's accuracy and response quality.
[0706] Through the processing steps described above, this system can mimic the knowledge and judgment of managers in real time, thereby improving the efficiency of business operations and the accuracy of decision-making.
[0707] (Application Example 1)
[0708] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0709] In recent years, there has been a growing demand for increased efficiency in corporate management and improved accuracy in decision-making. However, means of passing on leaders' knowledge and judgment have been limited, and providing real-time advice has been difficult. In particular, there is a need for improved quality of immediate information provision and advice to users in virtual environments. Furthermore, when using virtual reality devices, the lack of appropriate conversational interfaces to enhance the user experience is a challenge.
[0710] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0711] In this invention, the server includes means for collecting information, means for preprocessing the collected information to remove noise and unify the data format, means for training a generative AI model using the preprocessed data to generate a virtual mentor emulator, means for interacting with the virtual mentor through a user interface, means for collecting interaction data and continuously retraining the AI model, and means for providing the user with advice on product selection within a virtual environment through a virtual reality device. This enables increased efficiency in business management and improved accuracy in decision-making, and provides a real-time conversational interface that enhances the user experience in a virtual environment.
[0712] "Information" refers to a series of data related to the user, such as documents, data, audio, and images.
[0713] "Collection" refers to the entire process of gathering information.
[0714] "Preprocessing" refers to the process of removing noise from collected information and standardizing the data format.
[0715] A "generative AI model" refers to an algorithm or system that uses a large dataset for artificial intelligence to learn and perform a specified task.
[0716] A "virtual leader emulator" refers to a virtual agent created to mimic the speech patterns and behavioral patterns of a specific leader.
[0717] A "user interface" refers to the interface through which a user interacts with a system.
[0718] "Interaction data" refers to data related to the interactions and operations that take place between the user and the system.
[0719] "Retraining an AI model" refers to the process of retraining an existing AI model to improve its performance based on new data.
[0720] "Virtual reality devices" refer to hardware and software used to provide a virtual reality (VR) environment.
[0721] "Product selection advice" refers to information and suggestions provided by a supervisor emulator to help users choose a specific product.
[0722] "Real-time" refers to the process of generating an immediate response to user input.
[0723] System Overview
[0724] This invention realizes a system that generates a virtual instructor emulator and provides advice to users through a virtual environment or virtual reality device. Specifically, users can move freely within a virtual store and receive real-time advice from a virtual instructor regarding product selection.
[0725] Hardware and software used
[0726] Server: Performs data collection, preprocessing, training and retraining of AI models, and generation of virtual leader emulators.
[0727] Virtual reality devices (HMDs, etc.): Used by users to move around within a virtual store and interact with a virtual instructor emulator through an interface.
[0728] Generative AI Models: Using GPT-3 or similar high-performance AI models, the system generates statements and advice from a virtual leader emulator.
[0729] Natural language processing tools: Used for analyzing and preprocessing collected text data.
[0730] Processing flow and data calculations
[0731] 1. Information Gathering: The server uses web scraping techniques to collect publicly available information and internal documents. This includes social media posts, interview articles, and publicly available emails.
[0732] 2. Preprocessing: The collected information is preprocessed to remove noise and standardize the data format. Specifically, grammar checking tools are used to correct typographical errors and unnecessary HTML tags and special characters are removed.
[0733] 3. Training Generative AI Models: GPT-3 or similar generative AI models are trained using preprocessed data. During training, the data is divided into a training set and a test set, and the model is optimized.
[0734] 4. Generation of a virtual leader emulator: The server generates a virtual leader emulator using a trained AI model. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time.
[0735] 5. Collection and retraining of interaction data: Data generated when users interact with the instructor emulator using a virtual reality device is collected and used to continuously retrain the AI model.
[0736] Specific example
[0737] In a virtual store:
[0738] The user wears an HMD (Head-Mounted Display) and moves around within the virtual store.
[0739] The participant makes the gesture of picking up any product and is asked, "What are the features of this product?"
[0740] A virtual coach emulator provides real-time responses such as, "This product utilizes the latest running technology, is extremely lightweight, and durable. It is especially suitable for long-distance running."
[0741] Example of a prompt
[0742] User: "Could you please give me some advice about this new smartphone?"
[0743] Virtual Leader Emulator: "This smartphone features the latest processor, making it extremely fast and offering excellent battery life. It's especially recommended for gaming and video editing."
[0744] Thus, this invention allows users to receive real-time advice from a mentor emulator when effectively selecting products in a virtual environment. This aims to improve the efficiency of business operations and the accuracy of decision-making.
[0745] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0746] Step 1:
[0747] The server collects information. Specifically, it uses web scraping techniques to collect publicly available information from the web (such as social media posts, blog articles, interviews, and publicly available emails) and stores it in a database. The input is a URL or API endpoint on the web, and the output is the collected text data.
[0748] Step 2:
[0749] The server preprocesses the collected information. Specifically, it uses a grammar checker to correct typos and removes HTML tags and special characters. Next, it analyzes the text data using natural language processing techniques to extract important keywords and sentences. The input is the collected text data, and the output is text data in a clean and consistent format.
[0750] Step 3:
[0751] The server trains a generative AI model using preprocessed data. Specifically, it uses a model such as GPT-3, training the model by splitting the data into training and test sets. It optimizes the model parameters using backpropagation. The input is preprocessed text data, and the output is the trained AI model.
[0752] Step 4:
[0753] The server generates a virtual leader emulator using a trained AI model. This emulator has speech recognition and speech synthesis capabilities and is deployed on the cloud. The input is the trained AI model, and the output is the virtual leader emulator.
[0754] Step 5:
[0755] The user navigates a virtual environment through a virtual reality device (HMD) and selects products. They use an interface to interact with a virtual instructor emulator and ask questions about the products. Input is either voice or text input from the user, and output is a voice response from the virtual instructor emulator.
[0756] Step 6:
[0757] The server collects interaction data between the user and the virtual instructor emulator. Specifically, it collects data including dialogue history and usage patterns, and uses this data to retrain the AI model. The input is interaction data, and the output is the retrained AI model.
[0758] Step 7:
[0759] The server deploys a retrained AI model to the cloud, improving the accuracy and responsiveness of the virtual mentor emulator. This allows users to receive up-to-date information and optimal advice in real time. The input is the retrained AI model, and the output is the improved virtual mentor emulator.
[0760] By following these steps, a system will be completed that improves the efficiency of business operations and the accuracy of decision-making.
[0761] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0762] System Overview
[0763] This invention relates to a system that generates a virtual CEO emulator by collecting and preprocessing information related to the CEO, and training a generative AI model and emotion engine, enabling interaction through a user interface. This system can recognize the user's emotions and adjust its responses based on those emotions.
[0764] Program Processing Overview
[0765] This system consists of the following main phases:
[0766] 1. Data Collection Phase
[0767] 2. Data preprocessing phase
[0768] 3. Model training phase
[0769] 4. Emotional Engine Learning Phase
[0770] 5. CEO Emulator Generation Phase
[0771] 6. Interaction Phase
[0772] Program operation description
[0773] Data collection phase
[0774] The server uses web scraping techniques to collect data related to the CEO from specified websites and social media platforms. This includes publicly available blog posts, interviews, and social media posts by the CEO.
[0775] Users can use the system interface to upload additional CEO-related data, such as internal documents and past email data.
[0776] Data storage
[0777] The server stores the collected data in a database for centralized management. The stored data includes various formats such as text, images, and audio files.
[0778] Noise reduction and data cleansing
[0779] The server performs noise reduction processing on the collected data. This includes tasks such as correcting spelling errors and removing unnecessary tags and special characters.
[0780] The server standardizes data formats and cleans up text data. For example, it removes HTML tags and converts it to plain text.
[0781] Data analysis using natural language processing
[0782] The server uses natural language processing techniques to analyze clean text data. It performs tokenization, part-of-speech tagging, and sentence analysis to extract important keywords and sentences.
[0783] Creating a dataset
[0784] The server splits the analyzed data into a training set and a test set. For example, 70% of the data might be used for the training set and 30% for the test set.
[0785] Training of generative AI models
[0786] The server trains a generative AI model using a training set. It uses deep learning algorithms and performs backpropagation to optimize the model's parameters.
[0787] The server trains the model multiple times, specifying the number of epochs, to improve the model's accuracy.
[0788] Learning the Emotion Engine
[0789] The server trains the emotion engine using a training set. It analyzes user voice and text input data, assigns emotion labels, and trains the model.
[0790] The server will implement a function that uses an emotion engine to recognize emotions in real time from user input.
[0791] Model evaluation
[0792] The server evaluates the performance of the trained generative AI models and sentiment engines using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[0793] The server retrains the model as needed to further improve its performance.
[0794] Creating a CEO emulator
[0795] The server uses the final generative AI model and emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time.
[0796] User interface deployment
[0797] The server deploys the generated virtual CEO emulator to the cloud. It provides a URL that users can access via a browser, and functions as the user interface.
[0798] Start interaction with the user
[0799] Users initiate a conversation with a virtual CEO using a web interface. They can ask questions and seek advice using text or voice input.
[0800] The server analyzes user input in real time and generates an appropriate response. The generated response is then delivered to the user as speech output using a speech synthesis engine.
[0801] The server analyzes the user's emotions using an emotion engine and adjusts its response based on the recognized emotions. For example, if the user is angry, it will respond in a calm tone.
[0802] Collection of interaction data and retraining
[0803] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement.
[0804] The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue.
[0805] Specific example
[0806] Data collection examples
[0807] The server collects the CEO's social media posts and interview articles from the past five years and stores them in a database.
[0808] Users can supplement the collected data by uploading personal emails and internal documents to the system.
[0809] Data preprocessing example
[0810] The server analyzes the collected data and corrects typos and grammatical errors. It also removes unnecessary tags and special characters, converting the data into clean text.
[0811] Model Learning Example
[0812] The server uses pre-processed data to train generative AI models and emotion engines, improving their ability to understand the CEO's speaking patterns and users' emotions.
[0813] Emulator generation example
[0814] The server generates a 2D virtual CEO and implements speech recognition and synthesis capabilities. These emulators are deployed on the cloud as the user interface.
[0815] User interaction examples
[0816] Users ask questions about project progress to a virtual CEO and receive specific advice. The server analyzes the user's questions and emotions in real time to generate the most appropriate answers.
[0817] Based on the above, the present invention provides a system that allows CEOs to leverage their knowledge and judgment into the future. This system recognizes user emotions and adjusts responses based on those emotions, enabling more natural and effective dialogue. This, in turn, improves the efficiency of corporate operations and the accuracy of decision-making.
[0818] The following describes the processing flow.
[0819] Step 1: Data Collection
[0820] The server collects data related to the CEO from specified websites and social media platforms using web scraping techniques. Specifically, it uses APIs to obtain social media posts and crawling techniques to obtain text data from blog posts and interviews.
[0821] Users upload additional CEO-related data, such as internal company documents and past email data, using the system interface. For example, they can upload files using a drag-and-drop function.
[0822] Step 2: Save
[0823] The server stores the collected data in a database for centralized management. For example, it might use a NoSQL database to store data in different formats (text, images, audio files, etc.).
[0824] Step 3: Noise Reduction and Data Cleansing
[0825] The server performs noise reduction on the collected data. This is done using automated scripts and natural language processing tools to correct spelling errors, remove unnecessary tags and special characters, and so on.
[0826] The server standardizes data formats and cleans up text data. For example, it removes HTML tags and converts it to plain text.
[0827] Step 4: Data analysis using natural language processing
[0828] The server uses natural language processing techniques to analyze clean text data. Specifically, it performs processes such as tokenization, part-of-speech tagging, sentence analysis, and semantic analysis.
[0829] The server extracts important keywords and sentences from the analysis results and stores them in a database. For example, it may use text frequency analysis or TF-IDF scoring.
[0830] Step 5: Create a dataset
[0831] The server splits the analyzed data into a training set and a test set. For example, 70% of the data could be used for the training set and 30% for the test set. The script is then executed to distribute the data randomly.
[0832] Step 6: Training the Generative AI Model
[0833] The server trains a generative AI model (e.g., GPT-3 or BERT) using the training set. It then performs backpropagation to optimize the model's parameters using a deep learning algorithm.
[0834] The server trains the model multiple times, specifying the number of epochs (the number of training iterations), to improve the model's accuracy. Batch processing is used to perform calculations efficiently during this process.
[0835] Step 7: Learning the Emotion Engine
[0836] The server simultaneously trains the emotion engine using the training set. It analyzes user voice and text input data, assigns emotion labels, and trains the model. Specifically, it performs supervised learning using labeled datasets.
[0837] The server implements a function to recognize emotions in real time from user input using an emotion engine. For example, it uses speech feature extraction for emotion analysis from speech and emotion classification algorithms for emotion analysis from text.
[0838] Step 8: Model Evaluation
[0839] The server evaluates the performance of the trained generative AI models and sentiment engines using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[0840] The server retrains the model as needed to further improve its performance, for example, by performing hyperparameter tuning.
[0841] Step 9: Generate the CEO emulator
[0842] The server uses the final generative AI model and emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time. A graphical user interface (GUI) is designed to be user-friendly.
[0843] Step 10: Deploying the User Interface
[0844] The server deploys the generated virtual CEO emulator to the cloud. For example, it uses a web application framework to provide a URL that users can access through a browser, thus functioning as a user interface.
[0845] Step 11: Start interacting with the user
[0846] Users initiate a conversation with a virtual CEO using a web interface. They can ask questions and seek advice using text or voice input. The service operates in a browser and uses a microphone for voice input.
[0847] The server analyzes user input (text or voice) in real time and generates an appropriate response. The generated response is then delivered to the user as voice output using a speech synthesis engine. For example, natural language generation (NLG) technology is used to generate text, and a text-to-speech (TTS) engine is used to generate speech.
[0848] The server analyzes the user's emotions using an emotion engine and adjusts its response based on the recognized emotions. For example, if the user is angry, it will respond in a calm tone; if the user is sad, it will respond in a comforting tone.
[0849] Step 12: Collect interaction data and retrain
[0850] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement. The interaction logs stored in the database include user questions, virtual CEO responses, and user sentiment information.
[0851] The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue. Regular retraining allows the system to evolve over time, enabling more accurate responses.
[0852] (Example 2)
[0853] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0854] In current corporate operations, there is a lack of means to virtually replicate the knowledge and judgment of the CEO and to facilitate effective communication with employees and stakeholders. Furthermore, there is no system that can recognize user emotions in real time and adjust responses accordingly. This can potentially lead to a decrease in the efficiency of decision-making and problem-solving within companies.
[0855] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0856] In this invention, the server includes means for collecting information related to the CEO, means for preprocessing the collected information to remove noise and standardize the data format, means for analyzing the preprocessed data using natural language processing techniques to extract important keywords and sentences, means for training a generative AI model using the extracted data, means for generating a virtual CEO emulator with emotion recognition capabilities using the generated AI model, means for interacting with the virtual CEO through a user interface, means for collecting interaction data and continuously retraining the AI model, means for recognizing emotions in real time from user input, and means for adjusting responses based on recognized emotions. This makes it possible to virtually reproduce the knowledge and judgment of the CEO and generate appropriate responses according to the user's emotions.
[0857] "CEO-related information" refers to text, image, and audio data related to the CEO, such as blog posts, interviews, social media posts, internal documents, and past email data.
[0858] "Means of collection" refers to systems that include methods for obtaining specified information from the internet using web scraping technology, as well as methods for uploading internal documents and past email data via a user interface.
[0859] "Methods for preprocessing, noise reduction, and data format standardization" refer to techniques for processing collected data to create a clean data format by correcting spelling errors, removing unnecessary tags and special characters, and converting HTML tags to plain text.
[0860] "Natural language processing technology" refers to algorithms and methods used to analyze text data and extract important keywords and sentences through tokenization, part-of-speech tagging, and sentence analysis.
[0861] A "generative AI model" is an AI model trained using deep learning algorithms that has the ability to generate appropriate outputs for specific inputs.
[0862] A "virtual CEO emulator" is a virtual character created based on a generative AI model, possessing speech recognition and speech synthesis capabilities, and serving as an interface that can interact with users in real time.
[0863] A "user interface" refers to the graphical screen or web page that a user uses to interact with a virtual CEO.
[0864] "Interaction data" refers to digital information, including dialogue logs and response data from interactions between the user and the virtual CEO emulator.
[0865] "Emotion recognition functionality" refers to technology that analyzes user voice and text input data and identifies the user's emotions based on emotion labels extracted from that data.
[0866] "Means of adjusting responses based on emotions" refers to methods or algorithms for adjusting the content and tone of generated responses according to the user's emotions identified by emotion recognition functions.
[0867] This invention is a system that collects information related to a CEO, trains a generative AI model based on that information to generate a virtual CEO emulator, and allows the user to interact with it through an interface. This system also has the ability to recognize the user's emotions and adjust its responses based on those emotions. The specific embodiments of this system are described in detail below.
[0868] Data collection and storage
[0869] The server collects information related to the CEO from specified websites and social media platforms. This task utilizes web scraping techniques such as BeautifulSoup and Scrapy. The collected information includes the CEO's blog posts, interviews, and social media posts. The collected data is stored in databases such as MySQL and MongoDB.
[0870] Specific example:
[0871] The server collects social media posts and interview articles from the past five years and stores them in a database.
[0872] Users upload internal documents and past email data using the system interface.
[0873] Data preprocessing
[0874] The server performs noise reduction and data formatting on the collected data. Specifically, it corrects spelling errors using regular expressions and removes unnecessary HTML tags and special characters. Text data is converted to plain text. Python is often used for this process.
[0875] Analysis using natural language processing
[0876] The server applies natural language processing techniques to the pre-processed data. For example, NLTK or SpaCy are used to tokenize text data, tag parts of speech, and analyze sentences, extracting important keywords and sentences.
[0877] Specific example:
[0878] The server tokenizes the text data, tags it with parts of speech, and extracts important keywords.
[0879] Model training
[0880] The server trains generative AI models using the extracted data. Deep learning frameworks used include TensorFlow and PyTorch. Parameter optimization is performed using backpropagation with the training set.
[0881] Learning the Emotion Engine
[0882] The server trains an emotion engine, as well as a trained generative AI model. This emotion engine uses a BERT-based emotion analysis model to extract emotion labels (e.g., joy, sadness, anger) from the user's voice and text input data.
[0883] Specific example:
[0884] The server analyzes the audio data and trains a model to assign emotion labels.
[0885] Creating a virtual emulator
[0886] The server integrates a generative AI model and an emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition (Google Speech-to-Text API) and speech synthesis (Amazon Polly) capabilities.
[0887] Deployment and User Interaction
[0888] The server deploys the generated virtual CEO emulator to a cloud environment (AWS or Azure), making it accessible to users via a browser. Users can initiate an interaction with the virtual CEO using the interface. The server analyzes questions entered via text or voice in real time and generates appropriate responses. These responses are output as speech using a speech synthesis engine. It also features an emotion engine that analyzes the user's emotions and adjusts the tone of the responses accordingly.
[0889] Specific example:
[0890] The user enters, "Based on the CEO's latest social media posts, please share your insights into current market trends."
[0891] The server analyzes this input and generates a response through a virtual CEO emulator.
[0892] Collection of interaction data and retraining
[0893] The server collects user interaction logs and stores them as interaction data. Based on this data, the model is retrained to continuously improve performance.
[0894] The above describes a specific embodiment for implementing the present invention, which will lead to increased efficiency in decision-making and problem-solving within companies.
[0895] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0896] Step 1: Data Collection
[0897] The server collects information related to the CEO from specified websites and social media platforms. Specifically, it uses web scraping techniques (e.g., BeautifulSoup, Scrapy) to obtain the data.
[0898] Users upload internal documents and past email data through the system interface.
[0899] Input: Website URL, social media platform account information, user-uploaded internal documents and email data.
[0900] Output: Collected text data, image data, and audio data.
[0901] Step 2: Save Data
[0902] The server stores the collected data in a database (e.g., MySQL, MongoDB). Text data, image data, audio data, etc., are stored in the appropriate format according to their respective requirements.
[0903] Input: Collected text data, image data, and audio data.
[0904] Output: Clean data stored in the database.
[0905] Step 3: Noise Reduction and Data Cleansing
[0906] The server performs noise reduction processing on the collected data. Specifically, it uses regular expressions to correct spelling mistakes and remove unnecessary HTML tags and special characters.
[0907] Input: Raw data stored in the database.
[0908] Output: Denoised and cleaned text data.
[0909] Step 4: Data analysis using natural language processing
[0910] The server uses pre-processed data to apply natural language processing techniques (e.g., NLTK, SpaCy) to extract important keywords and sentences through tokenization, part-of-speech tagging, and sentence analysis.
[0911] Input: Clean text data.
[0912] Output: Tokenized keywords, part-of-speech tags, and parsed sentences.
[0913] Step 5: Create a dataset
[0914] The server splits the analyzed data into a training set (e.g., 70% of the data) and a test set (e.g., 30% of the data). These datasets are used for model training and evaluation.
[0915] Input: Analyzed keywords and sentences.
[0916] Output: Training set, test set.
[0917] Step 6: Training the Generative AI Model
[0918] The server trains a generative AI model (e.g., GPT-3) using a training set. Deep learning frameworks used include TensorFlow and PyTorch. Parameter optimization is performed using backpropagation.
[0919] Input: Training set.
[0920] Output: A trained generative AI model.
[0921] Step 7: Learning the Emotion Engine
[0922] The server trains an emotion engine (e.g., a BERT-based emotion analysis model) using a training set. The model is trained by assigning emotion labels to user voice and text input data.
[0923] Input: Training set, user voice and text data.
[0924] Output: Trained emotion engine.
[0925] Step 8: Model Evaluation
[0926] The server evaluates the performance of generative AI models and emotion engines trained using a test set. Evaluation metrics include accuracy, recall, and F1 score, and retraining is performed as needed.
[0927] Input: Test set.
[0928] Output: Evaluation results (accuracy, recall, F1 score).
[0929] Step 9: Generate the CEO emulator
[0930] The server integrates the final generative AI model and emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition (Google Speech-to-Text API) and speech synthesis (Amazon Polly) capabilities.
[0931] Input: Trained generative AI model, emotion engine.
[0932] Output: Virtual CEO emulator.
[0933] Step 10: Deploying the User Interface
[0934] The server deploys the generated virtual CEO emulator to a cloud environment (AWS or Azure) and makes it accessible to users through a browser.
[0935] Input: Virtual CEO emulator.
[0936] Output: Deployed interface (URL).
[0937] Step 11: Interacting with the user
[0938] Users initiate a conversation with a virtual CEO through their browser, entering questions and concerns via text or voice.
[0939] The server analyzes user input in real time and generates an appropriate response. The generated response is delivered to the user using a speech synthesis engine. Additionally, an emotion engine analyzes the user's emotions and adjusts the tone of the response accordingly.
[0940] Input: User text and voice input.
[0941] Output: Virtual CEO's response (text and audio).
[0942] Step 12: Collect interaction data and retrain
[0943] The server collects user interaction logs and stores them in a database as interaction data. The collected data is used to retrain the emotion engine and generative AI models.
[0944] Input: Dialogue log between the user and the virtual CEO.
[0945] Output: Collected dialogue logs, retrained AI model, and emotion engine.
[0946] (Application Example 2)
[0947] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0948] In modern factory operations, diverse data needs to be collected and analyzed in real time, but systems for handling this data efficiently and effectively are not yet widespread. Furthermore, there is a lack of means to provide appropriate instructions and advice in real time based on worker emotions and production status, which can lead to problems such as decreased production efficiency and increased worker stress.
[0949] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting information related to the CEO, means for preprocessing the collected information to remove noise and unify the data format, means for training a generative AI model using the preprocessed data to generate a virtual CEO emulator, means for interacting with the virtual CEO through a user interface, means for collecting interaction data and continuously retraining the AI model, means for collecting and monitoring data on each machine and worker in the factory in real time, and means for analyzing production status and worker emotions based on the collected data, and for the virtual CEO to provide appropriate instructions and advice. This makes it possible to improve the efficiency of factory operations and reduce worker stress.
[0950] "Information related to the CEO" refers to publicly available information, internal documents, past email data, blog posts, interviews, and social media posts concerning the company's chief executive officer.
[0951] "Preprocessing" refers to a series of processes to remove noise from collected data and standardize the data format.
[0952] "Noise reduction" refers to the process of removing unnecessary elements (such as spelling mistakes, special characters, and HTML tags) from collected data.
[0953] "Data format standardization" refers to the process of converting data in different formats into a consistent format.
[0954] A "generative AI model" refers to an artificial intelligence model that has been trained using pre-collected data and is capable of generating appropriate responses to specific tasks.
[0955] A "virtual CEO emulator" refers to a virtual character based on a generative AI model trained to mimic the speech patterns and decision-making styles of a CEO.
[0956] "User interface" refers to the interactive means by which a user interacts with a system, including browsers, displays, etc.
[0957] "Interaction data" refers to data that records the dialogue logs and responses between the user and the virtual CEO emulator.
[0958] "Continuously retraining an AI model" refers to the process of periodically retraining an AI model to improve its accuracy based on collected interaction data.
[0959] "Data on each machine and worker within the factory" refers to real-time data such as the operating status of factory production equipment, the behavior of workers, and their emotional states.
[0960] "Real-time data collection and monitoring" refers to the process of instantly acquiring current conditions and continuously monitoring that data.
[0961] "Production status" refers to information related to the progress of product production within the factory and the operating status of machinery.
[0962] "Worker emotions" refers to data related to emotions, such as workers' stress levels, satisfaction levels, and fatigue levels.
[0963] "Providing instructions and advice" refers to the process by which the virtual CEO emulator recommends specific actions and measures based on the data it has collected.
[0964] This invention relates to a system using a virtual CEO emulator aimed at improving the efficiency of factory operations and reducing worker stress. The system has the following configuration:
[0965] Hardware and software to be used
[0966] The server will use a high-performance server, natural language processing libraries (such as SpaCy and NLTK), deep learning frameworks (such as TensorFlow and PyTorch), sentiment recognition libraries (such as openSMILE), and web scraping tools (such as BeautifulSoup and Scrapy).
[0967] The devices include IoT sensors within the factory, wearable devices for workers, tablets, and displays.
[0968] Program processing flow
[0969] Data collection
[0970] The server collects data in real time from IoT sensors and wearable devices within the factory. It also periodically collects relevant industry news and technical literature using web scraping tools.
[0971] Specific example:
[0972] "Workers' smartwatches measure heart rate and stress levels and transmit the data to a cloud server in real time."
[0973] "IoT sensors monitor the operating status of each machine and transmit the data to a server."
[0974] Data preprocessing
[0975] The server removes noise from the collected data and standardizes the data format. Specifically, it corrects typos and removes unnecessary tags and special characters.
[0976] Model Learning
[0977] The server uses pre-processed data to train generative AI models and emotion engines. This allows a virtual CEO emulator to understand the emotional state and productivity of workers and provide appropriate instructions and advice.
[0978] Generating a virtual CEO emulator
[0979] The server generates a virtual CEO emulator using the generated generative AI model and emotion engine. This emulator can interact in real time with workers and managers via tablets or displays.
[0980] Example of a prompt:
[0981] "Could you tell me the current status of the production line?"
[0982] "Monitor the stress levels of the workers and suggest breaks if necessary."
[0983] User Interface and Interaction
[0984] Users can receive real-time feedback and advice on production status and individual issues through interaction with a virtual CEO emulator. The server has the capability to analyze user input and emotions and generate appropriate responses.
[0985] Specific example:
[0986] "Workers ask a virtual CEO questions about the project's progress and receive specific advice. The server analyzes the user's questions and emotions in real time to generate the most appropriate response."
[0987] Collection of interaction data and retraining
[0988] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement. The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue.
[0989] In this way, the present invention improves the efficiency of factory operations and reduces stress on workers.
[0990] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0991] Step 1:
[0992] The server collects data in real time from IoT sensors and wearable devices within the factory. Inputs include data such as the operating status of each machine, workers' heart rates, and stress levels. Based on this, the server centrally stores this data in a database, making it immediately accessible.
[0993] Step 2:
[0994] The server preprocesses the collected data. The input is the raw data collected in step 1. Specifically, it corrects typos and grammatical errors, removes HTML tags and unnecessary special characters, and standardizes the data into a clean text format. The output is the clean data saved in the standardized format.
[0995] Step 3:
[0996] The server applies natural language processing techniques to the pre-processed data to extract important keywords and sentences. The input is the clean text data from step 2. Specific operations include tokenization, part-of-speech tagging, and sentence analysis. The output is a list of the extracted keywords and sentences.
[0997] Step 4:
[0998] The server trains a generative AI model using preprocessed data and natural language processing results. The input is the data from steps 2 and 3. A deep learning framework (such as TensorFlow or PyTorch) is used, and backpropagation is performed to optimize the model parameters. The output is the trained generative AI model.
[0999] Step 5:
[1000] The server trains its emotion engine using an emotion recognition library (such as openSMILE). The input consists of worker voice and text data. Its specific actions include voice analysis and emotion labeling. The output is the trained emotion engine.
[1001] Step 6:
[1002] The server generates a virtual CEO emulator using a trained generative AI model and an emotion engine. The input is the trained model from steps 4 and 5. The virtual CEO emulator can interact with the user in real time via a tablet or display. The output is the virtual CEO emulator running on the user interface.
[1003] Step 7:
[1004] The user begins interacting with a virtual CEO emulator. Specifically, they input prompts such as, "Please tell me the current status of the production line," or "Please monitor the stress levels of the workers and suggest breaks if necessary." The input can be text or voice from the user. The server analyzes this and generates the most appropriate response. The output is appropriate advice or instructions provided to the user.
[1005] Step 8:
[1006] The server collects user interaction data and stores it in the interaction database. The input is the interaction log from step 7. The output is the stored interaction data.
[1007] Step 9:
[1008] The server retrains the generative AI model and emotion engine using the collected interaction data. The input is the interaction data from step 8. As a result of the retraining, the accuracy of the AI model and emotion engine improves, enabling more natural and effective dialogue. The output is the further improved generative AI model and emotion engine.
[1009] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1010] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1011] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1012] [Third Embodiment]
[1013] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1014] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1015] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1016] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1017] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1018] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1019] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1020] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1021] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1022] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1023] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1024] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1025] System Overview
[1026] This invention relates to a system that generates a virtual CEO emulator by collecting and preprocessing information related to the CEO and training a generative AI model, enabling interaction through a user interface. This system can learn the CEO's statements and decisions, imitate them in real time, and interact with the user via a two-dimensional interface.
[1027] Program Processing Overview
[1028] This system consists of the following main phases:
[1029] 1. Data Collection Phase
[1030] 2. Data preprocessing phase
[1031] 3. Model training phase
[1032] 4. CEO Emulator Generation Phase
[1033] 5. Interaction Phase
[1034] Program operation description
[1035] Data collection phase
[1036] The server collects publicly available information related to the CEO using web scraping techniques. This includes social media posts, blog articles, interviews, and publicly available emails.
[1037] Users can improve the accuracy of data collection by uploading internal documents and files to the system.
[1038] Data preprocessing phase
[1039] The server preprocesses the collected data, removing noise and standardizing the data format. Specifically, it uses a grammar checking tool to correct typos and removes unnecessary HTML tags and special characters.
[1040] The server uses natural language processing technology to analyze text data and extract important keywords and sentences.
[1041] Model Learning Phase
[1042] The server uses a generative AI model (e.g., GPT-3 or BERT) to train on preprocessed data. Initial setup is performed, and the model is instantiated.
[1043] The server splits the data into a training set and a test set and performs training. During this process, it optimizes the model parameters using backpropagation.
[1044] The server evaluates the model's performance using a test set and retrains it as needed.
[1045] CEO Emulator Generation Phase
[1046] The server generates a virtual CEO emulator as a 2D interface. This emulator is equipped with speech recognition and speech synthesis capabilities, enabling real-time interaction with the user.
[1047] The server deploys the generated interface as a user interface on the cloud, making it accessible to users.
[1048] Interaction Phase
[1049] Users interact with a virtual CEO through their device. For example, they log in to a web application using a browser and open a dialogue screen.
[1050] The server analyzes user input (text or voice) and generates an appropriate response in real time. The generated response is then delivered to the user as voice output through a speech synthesis engine.
[1051] The server continuously collects new dialogue data and retrains the model, thereby improving the system's accuracy and response quality.
[1052] Specific example
[1053] Data collection examples
[1054] The server collects the CEO's social media posts and interview articles from the past five years and stores them in a database.
[1055] Users can supplement the collected data by uploading personal emails and internal documents to the system.
[1056] Data preprocessing example
[1057] The server analyzes the collected data and corrects typos and grammatical errors. It also removes unnecessary tags and special characters, converting the data into clean text.
[1058] Model Learning Example
[1059] The server uses pre-processed data to train a generative AI model, learning the CEO's speaking patterns. For example, it incorporates frequently used phrases and decision-making tendencies into the model.
[1060] Emulator generation example
[1061] The server generates a 2D virtual CEO and implements speech recognition and synthesis capabilities. These emulators are deployed on the cloud as the user interface.
[1062] User interaction examples
[1063] Users ask questions about project progress to a virtual CEO and receive specific advice. The server analyzes the user's questions in real time and generates the most appropriate answers.
[1064] In summary, this invention provides a system that allows CEOs to leverage their knowledge and judgment for the future. This will lead to increased efficiency in corporate operations and improved accuracy in decision-making.
[1065] The following describes the processing flow.
[1066] Step 1: Data Collection
[1067] The server uses web scraping techniques to collect data related to the CEO from specified websites and social media platforms. This includes publicly available blog posts, interviews, and social media posts by the CEO.
[1068] Users can use the system interface to upload additional CEO-related data, such as internal documents and past email data.
[1069] Step 2: Save
[1070] The server stores the collected data in a database for centralized management. The stored data includes various formats such as text, images, and audio files.
[1071] Step 3: Noise Reduction and Data Cleansing
[1072] The server performs noise reduction processing on the collected data. This includes tasks such as correcting spelling errors and removing unnecessary tags and special characters.
[1073] The server standardizes data formats and cleans up text data. For example, it removes HTML tags and converts it to plain text.
[1074] Step 4: Data analysis using natural language processing
[1075] The server uses natural language processing techniques to analyze clean text data. It performs tokenization, part-of-speech tagging, and sentence analysis to extract important keywords and sentences.
[1076] Step 5: Create a dataset
[1077] The server splits the analyzed data into a training set and a test set. For example, 70% of the data might be used for the training set and 30% for the test set.
[1078] Step 6: Training the Generative AI Model
[1079] The server trains a generative AI model using a training set. It uses deep learning algorithms and performs backpropagation to optimize the model's parameters.
[1080] The server trains the model multiple times, specifying the number of epochs, to improve the model's accuracy.
[1081] Step 7: Model Evaluation
[1082] The server evaluates the performance of the trained model using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[1083] The server retrains the model as needed to further improve its performance.
[1084] Step 8: Generating the CEO Emulator
[1085] The server uses the final generative AI model to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time.
[1086] Step 9: Deploying the User Interface
[1087] The server deploys the generated virtual CEO emulator to the cloud. It provides a URL that users can access via a browser, and functions as the user interface.
[1088] Step 10: Start interacting with the user
[1089] Users initiate a conversation with a virtual CEO using a web interface. They can ask questions and seek advice using text or voice input.
[1090] The server analyzes user input in real time and generates an appropriate response. The generated response is then delivered to the user as speech output using a speech synthesis engine.
[1091] Step 11: Collect interaction data and retrain
[1092] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement.
[1093] The server uses the collected interaction data to retrain the model and improve the quality of the dialogue.
[1094] (Example 1)
[1095] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1096] In today's business environment, there is a growing need to quickly and accurately mimic the judgments and instructions of managers in order to improve the efficiency of corporate operations and the precision of decision-making. Furthermore, there is an increasing need for systems that can utilize the knowledge and experience of managers even when they are physically absent. However, conventional technologies do not adequately provide methods for mimicking managers' thought patterns and statements in real time.
[1097] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1098] In this invention, the server includes means for collecting publicly available information, means for preprocessing the collected information, performing grammar checks and unifying data formats, means for training a generative AI model using the preprocessed data to generate a virtual manager emulator, means for interacting with the virtual manager through a user interface, and means for collecting interaction data and continuously retraining the AI model. This enables real-time imitation of the manager's knowledge and judgment, improving the efficiency of business operations and the accuracy of decision-making.
[1099] "Public information" refers to information that is accessible to a large number of people, such as information posted on websites, social media posts, blog articles, interview articles, and publicly available emails.
[1100] "Preprocessing" is the process of correcting typographical errors in collected data, removing unnecessary HTML tags and special characters, and standardizing data formats.
[1101] A "generative AI model" is an artificial intelligence model that is trained using collected and pre-processed data, and includes, for example, natural language generation models such as GPT-3.
[1102] A "virtual executive emulator" is a virtual character that mimics the statements and decisions of executives in real time based on a generated AI model, and it has speech recognition and speech synthesis capabilities.
[1103] "User interface" refers to the screens and applications that allow a user to interact with a virtual manager emulator, and includes web browsers and dedicated applications.
[1104] "Interaction data" refers to the data from conversations between the user and a virtual manager emulator, and this data is used to retrain the AI model.
[1105] "Natural language processing technology" refers to techniques for analyzing text data and extracting important keywords and sentences, and includes libraries such as NLTK and SpaCy.
[1106] "Speech recognition" is a technology that converts speech data into text data, and Google Speech-to-Text is an example of this.
[1107] "Speech synthesis" is a technology that converts text data into speech data, and Amazon Polly is an example of this.
[1108] This invention is a system that generates a virtual CEO emulator by collecting and pre-processing information related to the CEO and training a generated AI model, enabling interaction through a user interface. This system imitates the judgments and statements of executives in real time based on publicly available information, thereby improving the efficiency of corporate operations and the accuracy of decision-making.
[1109] The system configuration is mainly described in the following phases:
[1110] 1. Data Collection Phase
[1111] 2. Data preprocessing phase
[1112] 3. Model training phase
[1113] 4. CEO Emulator Generation Phase
[1114] 5. Interaction Phase
[1115] Data collection phase
[1116] The server uses web scraping techniques (e.g., Beautiful Soup) to collect publicly available information such as social media posts, blog articles, interviews, and published emails. Specifically, it uses Python libraries to crawl websites containing certain keywords (e.g., names of business owners, company names).
[1117] Users can improve the quality and quantity of data collection by uploading internal documents and personal emails to the system. The uploaded files are stored in a database on the server.
[1118] Data preprocessing phase
[1119] The server corrects the collected data using a grammar checking tool (e.g., Grammarly API) to eliminate typos and grammatical errors.
[1120] The server removes unnecessary HTML tags and special characters from the collected data, converting it into clean text data.
[1121] The server analyzes text data using natural language processing technologies (e.g., NLTK, SpaCy) to extract important keywords and sentences. For example, it can extract key topics and keywords from interview articles with business executives.
[1122] Model Learning Phase
[1123] The server instantiates a generative AI model (e.g., GPT-3) and loads pre-processed data. Instances of the generative AI model are created using the OpenAI API.
[1124] The server splits the preprocessed data into a training set and a test set, and then trains the model.
[1125] The server uses backpropagation to optimize the parameters of the generated AI model and retrains it to improve its accuracy.
[1126] CEO Emulator Generation Phase
[1127] The server generates a virtual manager emulator using a pre-trained generative AI model. This emulator is designed as a two-dimensional interface and includes speech recognition (e.g., Google Speech-to-Text) and speech synthesis (e.g., Amazon Polly) capabilities.
[1128] The server deploys a virtual manager emulator on the cloud and makes it accessible to users. For example, it might be deployed on AWS and configured to be accessible to users via a browser.
[1129] Interaction Phase
[1130] The user logs into the web application from their device and opens the interactive interface.
[1131] Example: A user logs into the system using a browser and asks a virtual manager, "What is the current project status?"
[1132] The server analyzes user input (text or voice) and generates an appropriate response.
[1133] Example: The server analyzes the user's question and provides the generated response as speech output using a speech synthesis engine.
[1134] The server continuously collects new interaction data and retrains the generative AI model. This improves the system's accuracy and response quality.
[1135] Examples of specific prompt statements include the following:
[1136] "Could you tell me the CEO's view on the progress of the current project?"
[1137] "We'd like to hear the CEO's opinion on our new marketing strategy."
[1138] As described above, the present invention makes it possible to mimic the knowledge and judgment of managers in real time, thereby improving the efficiency of corporate operations and the accuracy of decision-making.
[1139] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1140] The specific processing flow of the program
[1141] Step 1: Data Collection
[1142] The server launches a web scraping script to collect publicly available information such as social media posts, blog articles, interviews, and published emails. Specifically, it uses the Python Beautiful Soup library to crawl websites containing specific keywords.
[1143] Input: Website URL or specific keywords
[1144] Output: Collected text data (e.g., text from social media posts or blog articles)
[1145] Specific operation: The server accesses a specific website, retrieves the page content, and extracts the necessary parts.
[1146] Step 2: Data Upload
[1147] Users upload internal documents and personal emails to the system. The uploaded files are stored in a database on the server.
[1148] Input: A file selected by the user (e.g., PDF, text file)
[1149] Output: Text data stored in the database
[1150] Specific operation: The user sends a file to the server using the system's upload form, and the server saves its contents to the database.
[1151] Step 3: Data Cleaning
[1152] The server uses a grammar checking tool (e.g., Grammarly API) to correct typos and grammatical errors in the collected data.
[1153] Input: Collected text data
[1154] Output: Clean text data with typos and grammatical errors corrected.
[1155] Specific operation: The server sends text data to a grammar checking tool and receives the corrected text.
[1156] Step 4: Unify text formatting
[1157] The server removes unnecessary HTML tags and special characters from the collected data to standardize the data format.
[1158] Input: Text data with typos and grammatical errors corrected.
[1159] Output: Clean text data
[1160] Specific operation: The server uses regular expressions to remove HTML tags and special characters, converting the data into clean text.
[1161] Step 5: Natural Language Processing (NLP)
[1162] The server uses natural language processing techniques (e.g., NLTK, SpaCy) to analyze text data and extract important keywords and sentences.
[1163] Input: Clean text data
[1164] Output: Extracted keywords and important sentences
[1165] Specific operation: The server starts a natural language processing library, analyzes the text data, and extracts important information.
[1166] Step 6: Model instantiation and data loading
[1167] The server instantiates a generative AI model (e.g., GPT-3) and loads pre-processed data.
[1168] Input: Extracted keywords and important sentences
[1169] Output: Text data loaded as training data
[1170] Specific operation: The server uses the OpenAI API to create a GPT-3 instance and load the pre-processed data.
[1171] Step 7: Model Training
[1172] The server divides the pre-processed data into a training set and a test set, and then trains the generative AI model.
[1173] Input: Loaded training data
[1174] Output: Trained generative AI model
[1175] Specific operation: The server optimizes the model using training data and evaluates its accuracy using test data.
[1176] Step 8: Generate CEO emulator
[1177] The server generates a virtual manager emulator using a pre-trained generative AI model. This emulator has a two-dimensional interface and features speech recognition and speech synthesis capabilities.
[1178] Input: Trained generative AI model
[1179] Output: Virtual manager emulator
[1180] Specific operation: The server generates characters from the model and integrates speech recognition and speech synthesis engines.
[1181] Step 9: Interface Deployment
[1182] The server deploys a virtual manager emulator on the cloud, making it accessible to users.
[1183] Input: Virtual manager emulator
[1184] Output: Deployed user interface
[1185] Specific operation: The server deploys the emulator to a cloud service such as AWS and configures it so that users can access it via a browser.
[1186] Step 10: User Interaction
[1187] The user logs into the web application from their device, opens a conversational interface, and interacts with a virtual manager.
[1188] Input: User voice or text input
[1189] Output: Response from the virtual manager
[1190] Specific operation: The user inputs a question or instruction, and the server generates a response and sends it back.
[1191] Step 11: Data Collection and Retraining
[1192] The server continuously collects conversational data between the user and the virtual manager, and uses it to retrain the generated AI model.
[1193] Input: User interaction data
[1194] Output: Retrained generative AI model
[1195] Specific operation: The server collects interaction data, periodically retrains the model, and improves the system's accuracy and response quality.
[1196] Through the processing steps described above, this system can mimic the knowledge and judgment of managers in real time, thereby improving the efficiency of business operations and the accuracy of decision-making.
[1197] (Application Example 1)
[1198] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1199] In recent years, there has been a growing demand for increased efficiency in corporate management and improved accuracy in decision-making. However, means of passing on leaders' knowledge and judgment have been limited, and providing real-time advice has been difficult. In particular, there is a need for improved quality of immediate information provision and advice to users in virtual environments. Furthermore, when using virtual reality devices, the lack of appropriate conversational interfaces to enhance the user experience is a challenge.
[1200] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1201] In this invention, the server includes means for collecting information, means for preprocessing the collected information to remove noise and unify the data format, means for training a generative AI model using the preprocessed data to generate a virtual mentor emulator, means for interacting with the virtual mentor through a user interface, means for collecting interaction data and continuously retraining the AI model, and means for providing the user with advice on product selection within a virtual environment through a virtual reality device. This enables increased efficiency in business management and improved accuracy in decision-making, and provides a real-time conversational interface that enhances the user experience in a virtual environment.
[1202] "Information" refers to a series of data related to the user, such as documents, data, audio, and images.
[1203] "Collection" refers to the entire process of gathering information.
[1204] "Preprocessing" refers to the process of removing noise from collected information and standardizing the data format.
[1205] A "generative AI model" refers to an algorithm or system that uses a large dataset for artificial intelligence to learn and perform a specified task.
[1206] A "virtual leader emulator" refers to a virtual agent created to mimic the speech patterns and behavioral patterns of a specific leader.
[1207] A "user interface" refers to the interface through which a user interacts with a system.
[1208] "Interaction data" refers to data related to the interactions and operations that take place between the user and the system.
[1209] "Retraining an AI model" refers to the process of retraining an existing AI model to improve its performance based on new data.
[1210] "Virtual reality devices" refer to hardware and software used to provide a virtual reality (VR) environment.
[1211] "Product selection advice" refers to information and suggestions provided by a supervisor emulator to help users choose a specific product.
[1212] "Real-time" refers to the process of generating an immediate response to user input.
[1213] System Overview
[1214] This invention realizes a system that generates a virtual instructor emulator and provides advice to users through a virtual environment or virtual reality device. Specifically, users can move freely within a virtual store and receive real-time advice from a virtual instructor regarding product selection.
[1215] Hardware and software used
[1216] Server: Performs data collection, preprocessing, training and retraining of AI models, and generation of virtual leader emulators.
[1217] Virtual reality devices (HMDs, etc.): Used by users to move around within a virtual store and interact with a virtual instructor emulator through an interface.
[1218] Generative AI Models: Using GPT-3 or similar high-performance AI models, the system generates statements and advice from a virtual leader emulator.
[1219] Natural language processing tools: Used for analyzing and preprocessing collected text data.
[1220] Processing flow and data calculations
[1221] 1. Information Gathering: The server uses web scraping techniques to collect publicly available information and internal documents. This includes social media posts, interview articles, and publicly available emails.
[1222] 2. Preprocessing: The collected information is preprocessed to remove noise and standardize the data format. Specifically, grammar checking tools are used to correct typographical errors and unnecessary HTML tags and special characters are removed.
[1223] 3. Training Generative AI Models: GPT-3 or similar generative AI models are trained using preprocessed data. During training, the data is divided into a training set and a test set, and the model is optimized.
[1224] 4. Generation of a virtual leader emulator: The server generates a virtual leader emulator using a trained AI model. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time.
[1225] 5. Collection and retraining of interaction data: Data generated when users interact with the instructor emulator using a virtual reality device is collected and used to continuously retrain the AI model.
[1226] Specific example
[1227] In a virtual store:
[1228] The user wears an HMD (Head-Mounted Display) and moves around within the virtual store.
[1229] The participant makes the gesture of picking up any product and is asked, "What are the features of this product?"
[1230] A virtual coach emulator provides real-time responses such as, "This product utilizes the latest running technology, is extremely lightweight, and durable. It is especially suitable for long-distance running."
[1231] Example of a prompt
[1232] User: "Could you please give me some advice about this new smartphone?"
[1233] Virtual Leader Emulator: "This smartphone features the latest processor, making it extremely fast and offering excellent battery life. It's especially recommended for gaming and video editing."
[1234] Thus, this invention allows users to receive real-time advice from a mentor emulator when effectively selecting products in a virtual environment. This aims to improve the efficiency of business operations and the accuracy of decision-making.
[1235] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1236] Step 1:
[1237] The server collects information. Specifically, it uses web scraping techniques to collect publicly available information from the web (such as social media posts, blog articles, interviews, and publicly available emails) and stores it in a database. The input is a URL or API endpoint on the web, and the output is the collected text data.
[1238] Step 2:
[1239] The server preprocesses the collected information. Specifically, it uses a grammar checker to correct typos and removes HTML tags and special characters. Next, it analyzes the text data using natural language processing techniques to extract important keywords and sentences. The input is the collected text data, and the output is text data in a clean and consistent format.
[1240] Step 3:
[1241] The server trains a generative AI model using preprocessed data. Specifically, it uses a model such as GPT-3, training the model by splitting the data into training and test sets. It optimizes the model parameters using backpropagation. The input is preprocessed text data, and the output is the trained AI model.
[1242] Step 4:
[1243] The server generates a virtual leader emulator using a trained AI model. This emulator has speech recognition and speech synthesis capabilities and is deployed on the cloud. The input is the trained AI model, and the output is the virtual leader emulator.
[1244] Step 5:
[1245] The user navigates a virtual environment through a virtual reality device (HMD) and selects products. They use an interface to interact with a virtual instructor emulator and ask questions about the products. Input is either voice or text input from the user, and output is a voice response from the virtual instructor emulator.
[1246] Step 6:
[1247] The server collects interaction data between the user and the virtual instructor emulator. Specifically, it collects data including dialogue history and usage patterns, and uses this data to retrain the AI model. The input is interaction data, and the output is the retrained AI model.
[1248] Step 7:
[1249] The server deploys a retrained AI model to the cloud, improving the accuracy and responsiveness of the virtual mentor emulator. This allows users to receive up-to-date information and optimal advice in real time. The input is the retrained AI model, and the output is the improved virtual mentor emulator.
[1250] By following these steps, a system will be completed that improves the efficiency of business operations and the accuracy of decision-making.
[1251] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1252] System Overview
[1253] This invention relates to a system that generates a virtual CEO emulator by collecting and preprocessing information related to the CEO, and training a generative AI model and emotion engine, enabling interaction through a user interface. This system can recognize the user's emotions and adjust its responses based on those emotions.
[1254] Program Processing Overview
[1255] This system consists of the following main phases:
[1256] 1. Data Collection Phase
[1257] 2. Data preprocessing phase
[1258] 3. Model training phase
[1259] 4. Emotional Engine Learning Phase
[1260] 5. CEO Emulator Generation Phase
[1261] 6. Interaction Phase
[1262] Program operation description
[1263] Data collection phase
[1264] The server uses web scraping techniques to collect data related to the CEO from specified websites and social media platforms. This includes publicly available blog posts, interviews, and social media posts by the CEO.
[1265] Users can use the system interface to upload additional CEO-related data, such as internal documents and past email data.
[1266] Data storage
[1267] The server stores the collected data in a database for centralized management. The stored data includes various formats such as text, images, and audio files.
[1268] Noise reduction and data cleansing
[1269] The server performs noise reduction processing on the collected data. This includes tasks such as correcting spelling errors and removing unnecessary tags and special characters.
[1270] The server standardizes data formats and cleans up text data. For example, it removes HTML tags and converts it to plain text.
[1271] Data analysis using natural language processing
[1272] The server uses natural language processing techniques to analyze clean text data. It performs tokenization, part-of-speech tagging, and sentence analysis to extract important keywords and sentences.
[1273] Creating a dataset
[1274] The server splits the analyzed data into a training set and a test set. For example, 70% of the data might be used for the training set and 30% for the test set.
[1275] Training of generative AI models
[1276] The server trains a generative AI model using a training set. It uses deep learning algorithms and performs backpropagation to optimize the model's parameters.
[1277] The server trains the model multiple times, specifying the number of epochs, to improve the model's accuracy.
[1278] Learning the Emotion Engine
[1279] The server trains the emotion engine using a training set. It analyzes user voice and text input data, assigns emotion labels, and trains the model.
[1280] The server will implement a function that uses an emotion engine to recognize emotions in real time from user input.
[1281] Model evaluation
[1282] The server evaluates the performance of the trained generative AI models and sentiment engines using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[1283] The server retrains the model as needed to further improve its performance.
[1284] Creating a CEO emulator
[1285] The server uses the final generative AI model and emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time.
[1286] User interface deployment
[1287] The server deploys the generated virtual CEO emulator to the cloud. It provides a URL that users can access via a browser, and functions as the user interface.
[1288] Start interaction with the user
[1289] Users initiate a conversation with a virtual CEO using a web interface. They can ask questions and seek advice using text or voice input.
[1290] The server analyzes user input in real time and generates an appropriate response. The generated response is then delivered to the user as speech output using a speech synthesis engine.
[1291] The server analyzes the user's emotions using an emotion engine and adjusts its response based on the recognized emotions. For example, if the user is angry, it will respond in a calm tone.
[1292] Collection of interaction data and retraining
[1293] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement.
[1294] The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue.
[1295] Specific example
[1296] Data collection examples
[1297] The server collects the CEO's social media posts and interview articles from the past five years and stores them in a database.
[1298] Users can supplement the collected data by uploading personal emails and internal documents to the system.
[1299] Data preprocessing example
[1300] The server analyzes the collected data and corrects typos and grammatical errors. It also removes unnecessary tags and special characters, converting the data into clean text.
[1301] Model Learning Example
[1302] The server uses pre-processed data to train generative AI models and emotion engines, improving their ability to understand the CEO's speaking patterns and users' emotions.
[1303] Emulator generation example
[1304] The server generates a 2D virtual CEO and implements speech recognition and synthesis capabilities. These emulators are deployed on the cloud as the user interface.
[1305] User interaction examples
[1306] Users ask questions about project progress to a virtual CEO and receive specific advice. The server analyzes the user's questions and emotions in real time to generate the most appropriate answers.
[1307] Based on the above, the present invention provides a system that allows CEOs to leverage their knowledge and judgment into the future. This system recognizes user emotions and adjusts responses based on those emotions, enabling more natural and effective dialogue. This, in turn, improves the efficiency of corporate operations and the accuracy of decision-making.
[1308] The following describes the processing flow.
[1309] Step 1: Data Collection
[1310] The server collects data related to the CEO from specified websites and social media platforms using web scraping techniques. Specifically, it uses APIs to obtain social media posts and crawling techniques to obtain text data from blog posts and interviews.
[1311] Users upload additional CEO-related data, such as internal company documents and past email data, using the system interface. For example, they can upload files using a drag-and-drop function.
[1312] Step 2: Save
[1313] The server stores the collected data in a database for centralized management. For example, it might use a NoSQL database to store data in different formats (text, images, audio files, etc.).
[1314] Step 3: Noise Reduction and Data Cleansing
[1315] The server performs noise reduction on the collected data. This is done using automated scripts and natural language processing tools to correct spelling errors, remove unnecessary tags and special characters, and so on.
[1316] The server standardizes data formats and cleans up text data. For example, it removes HTML tags and converts it to plain text.
[1317] Step 4: Data analysis using natural language processing
[1318] The server uses natural language processing techniques to analyze clean text data. Specifically, it performs processes such as tokenization, part-of-speech tagging, sentence analysis, and semantic analysis.
[1319] The server extracts important keywords and sentences from the analysis results and stores them in a database. For example, it may use text frequency analysis or TF-IDF scoring.
[1320] Step 5: Create a dataset
[1321] The server splits the analyzed data into a training set and a test set. For example, 70% of the data could be used for the training set and 30% for the test set. The script is then executed to distribute the data randomly.
[1322] Step 6: Training the Generative AI Model
[1323] The server trains a generative AI model (e.g., GPT-3 or BERT) using the training set. It then performs backpropagation to optimize the model's parameters using a deep learning algorithm.
[1324] The server trains the model multiple times, specifying the number of epochs (the number of training iterations), to improve the model's accuracy. Batch processing is used to perform calculations efficiently during this process.
[1325] Step 7: Learning the Emotion Engine
[1326] The server simultaneously trains the emotion engine using the training set. It analyzes user voice and text input data, assigns emotion labels, and trains the model. Specifically, it performs supervised learning using labeled datasets.
[1327] The server implements a function to recognize emotions in real time from user input using an emotion engine. For example, it uses speech feature extraction for emotion analysis from speech and emotion classification algorithms for emotion analysis from text.
[1328] Step 8: Model Evaluation
[1329] The server evaluates the performance of the trained generative AI models and sentiment engines using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[1330] The server retrains the model as needed to further improve its performance, for example, by performing hyperparameter tuning.
[1331] Step 9: Generate the CEO emulator
[1332] The server uses the final generative AI model and emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time. A graphical user interface (GUI) is designed to be user-friendly.
[1333] Step 10: Deploying the User Interface
[1334] The server deploys the generated virtual CEO emulator to the cloud. For example, it uses a web application framework to provide a URL that users can access through a browser, thus functioning as a user interface.
[1335] Step 11: Start interacting with the user
[1336] Users initiate a conversation with a virtual CEO using a web interface. They can ask questions and seek advice using text or voice input. The service operates in a browser and uses a microphone for voice input.
[1337] The server analyzes user input (text or voice) in real time and generates an appropriate response. The generated response is then delivered to the user as voice output using a speech synthesis engine. For example, natural language generation (NLG) technology is used to generate text, and a text-to-speech (TTS) engine is used to generate speech.
[1338] The server analyzes the user's emotions using an emotion engine and adjusts its response based on the recognized emotions. For example, if the user is angry, it will respond in a calm tone; if the user is sad, it will respond in a comforting tone.
[1339] Step 12: Collect interaction data and retrain
[1340] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement. The interaction logs stored in the database include user questions, virtual CEO responses, and user sentiment information.
[1341] The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue. Regular retraining allows the system to evolve over time, enabling more accurate responses.
[1342] (Example 2)
[1343] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1344] In current corporate operations, there is a lack of means to virtually replicate the knowledge and judgment of the CEO and to facilitate effective communication with employees and stakeholders. Furthermore, there is no system that can recognize user emotions in real time and adjust responses accordingly. This can potentially lead to a decrease in the efficiency of decision-making and problem-solving within companies.
[1345] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1346] In this invention, the server includes means for collecting information related to the CEO, means for preprocessing the collected information to remove noise and standardize the data format, means for analyzing the preprocessed data using natural language processing techniques to extract important keywords and sentences, means for training a generative AI model using the extracted data, means for generating a virtual CEO emulator with emotion recognition capabilities using the generated AI model, means for interacting with the virtual CEO through a user interface, means for collecting interaction data and continuously retraining the AI model, means for recognizing emotions in real time from user input, and means for adjusting responses based on recognized emotions. This makes it possible to virtually reproduce the knowledge and judgment of the CEO and generate appropriate responses according to the user's emotions.
[1347] "CEO-related information" refers to text, image, and audio data related to the CEO, such as blog posts, interviews, social media posts, internal documents, and past email data.
[1348] "Means of collection" refers to systems that include methods for obtaining specified information from the internet using web scraping technology, as well as methods for uploading internal documents and past email data via a user interface.
[1349] "Methods for preprocessing, noise reduction, and data format standardization" refer to techniques for processing collected data to create a clean data format by correcting spelling errors, removing unnecessary tags and special characters, and converting HTML tags to plain text.
[1350] "Natural language processing technology" refers to algorithms and methods used to analyze text data and extract important keywords and sentences through tokenization, part-of-speech tagging, and sentence analysis.
[1351] A "generative AI model" is an AI model trained using deep learning algorithms that has the ability to generate appropriate outputs for specific inputs.
[1352] A "virtual CEO emulator" is a virtual character created based on a generative AI model, possessing speech recognition and speech synthesis capabilities, and serving as an interface that can interact with users in real time.
[1353] A "user interface" refers to the graphical screen or web page that a user uses to interact with a virtual CEO.
[1354] "Interaction data" refers to digital information, including dialogue logs and response data from interactions between the user and the virtual CEO emulator.
[1355] "Emotion recognition functionality" refers to technology that analyzes user voice and text input data and identifies the user's emotions based on emotion labels extracted from that data.
[1356] "Means of adjusting responses based on emotions" refers to methods or algorithms for adjusting the content and tone of generated responses according to the user's emotions identified by emotion recognition functions.
[1357] This invention is a system that collects information related to a CEO, trains a generative AI model based on that information to generate a virtual CEO emulator, and allows the user to interact with it through an interface. This system also has the ability to recognize the user's emotions and adjust its responses based on those emotions. The specific embodiments of this system are described in detail below.
[1358] Data collection and storage
[1359] The server collects information related to the CEO from specified websites and social media platforms. This task utilizes web scraping techniques such as BeautifulSoup and Scrapy. The collected information includes the CEO's blog posts, interviews, and social media posts. The collected data is stored in databases such as MySQL and MongoDB.
[1360] Specific example:
[1361] The server collects social media posts and interview articles from the past five years and stores them in a database.
[1362] Users upload internal documents and past email data using the system interface.
[1363] Data preprocessing
[1364] The server performs noise reduction and data formatting on the collected data. Specifically, it corrects spelling errors using regular expressions and removes unnecessary HTML tags and special characters. Text data is converted to plain text. Python is often used for this process.
[1365] Analysis using natural language processing
[1366] The server applies natural language processing techniques to the pre-processed data. For example, NLTK or SpaCy are used to tokenize text data, tag parts of speech, and analyze sentences, extracting important keywords and sentences.
[1367] Specific example:
[1368] The server tokenizes the text data, tags it with parts of speech, and extracts important keywords.
[1369] Model training
[1370] The server trains generative AI models using the extracted data. Deep learning frameworks used include TensorFlow and PyTorch. Parameter optimization is performed using backpropagation with the training set.
[1371] Learning the Emotion Engine
[1372] The server trains an emotion engine, as well as a trained generative AI model. This emotion engine uses a BERT-based emotion analysis model to extract emotion labels (e.g., joy, sadness, anger) from the user's voice and text input data.
[1373] Specific example:
[1374] The server analyzes the audio data and trains a model to assign emotion labels.
[1375] Creating a virtual emulator
[1376] The server integrates a generative AI model and an emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition (Google Speech-to-Text API) and speech synthesis (Amazon Polly) capabilities.
[1377] Deployment and User Interaction
[1378] The server deploys the generated virtual CEO emulator to a cloud environment (AWS or Azure), making it accessible to users via a browser. Users can initiate an interaction with the virtual CEO using the interface. The server analyzes questions entered via text or voice in real time and generates appropriate responses. These responses are output as speech using a speech synthesis engine. It also features an emotion engine that analyzes the user's emotions and adjusts the tone of the responses accordingly.
[1379] Specific example:
[1380] The user enters, "Based on the CEO's latest social media posts, please share your insights into current market trends."
[1381] The server analyzes this input and generates a response through a virtual CEO emulator.
[1382] Collection of interaction data and retraining
[1383] The server collects user interaction logs and stores them as interaction data. Based on this data, the model is retrained to continuously improve performance.
[1384] The above describes a specific embodiment for implementing the present invention, which will lead to increased efficiency in decision-making and problem-solving within companies.
[1385] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1386] Step 1: Data Collection
[1387] The server collects information related to the CEO from specified websites and social media platforms. Specifically, it uses web scraping techniques (e.g., BeautifulSoup, Scrapy) to obtain the data.
[1388] Users upload internal documents and past email data through the system interface.
[1389] Input: Website URL, social media platform account information, user-uploaded internal documents and email data.
[1390] Output: Collected text data, image data, and audio data.
[1391] Step 2: Save Data
[1392] The server stores the collected data in a database (e.g., MySQL, MongoDB). Text data, image data, audio data, etc., are stored in the appropriate format according to their respective requirements.
[1393] Input: Collected text data, image data, and audio data.
[1394] Output: Clean data stored in the database.
[1395] Step 3: Noise Reduction and Data Cleansing
[1396] The server performs noise reduction processing on the collected data. Specifically, it uses regular expressions to correct spelling mistakes and remove unnecessary HTML tags and special characters.
[1397] Input: Raw data stored in the database.
[1398] Output: Denoised and cleaned text data.
[1399] Step 4: Data analysis using natural language processing
[1400] The server uses pre-processed data to apply natural language processing techniques (e.g., NLTK, SpaCy) to extract important keywords and sentences through tokenization, part-of-speech tagging, and sentence analysis.
[1401] Input: Clean text data.
[1402] Output: Tokenized keywords, part-of-speech tags, and parsed sentences.
[1403] Step 5: Create a dataset
[1404] The server splits the analyzed data into a training set (e.g., 70% of the data) and a test set (e.g., 30% of the data). These datasets are used for model training and evaluation.
[1405] Input: Analyzed keywords and sentences.
[1406] Output: Training set, test set.
[1407] Step 6: Training the Generative AI Model
[1408] The server trains a generative AI model (e.g., GPT-3) using a training set. Deep learning frameworks used include TensorFlow and PyTorch. Parameter optimization is performed using backpropagation.
[1409] Input: Training set.
[1410] Output: A trained generative AI model.
[1411] Step 7: Learning the Emotion Engine
[1412] The server trains an emotion engine (e.g., a BERT-based emotion analysis model) using a training set. The model is trained by assigning emotion labels to user voice and text input data.
[1413] Input: Training set, user voice and text data.
[1414] Output: Trained emotion engine.
[1415] Step 8: Model Evaluation
[1416] The server evaluates the performance of generative AI models and emotion engines trained using a test set. Evaluation metrics include accuracy, recall, and F1 score, and retraining is performed as needed.
[1417] Input: Test set.
[1418] Output: Evaluation results (accuracy, recall, F1 score).
[1419] Step 9: Generate the CEO emulator
[1420] The server integrates the final generative AI model and emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition (Google Speech-to-Text API) and speech synthesis (Amazon Polly) capabilities.
[1421] Input: Trained generative AI model, emotion engine.
[1422] Output: Virtual CEO emulator.
[1423] Step 10: Deploying the User Interface
[1424] The server deploys the generated virtual CEO emulator to a cloud environment (AWS or Azure) and makes it accessible to users through a browser.
[1425] Input: Virtual CEO emulator.
[1426] Output: Deployed interface (URL).
[1427] Step 11: Interacting with the user
[1428] Users initiate a conversation with a virtual CEO through their browser, entering questions and concerns via text or voice.
[1429] The server analyzes user input in real time and generates an appropriate response. The generated response is delivered to the user using a speech synthesis engine. Additionally, an emotion engine analyzes the user's emotions and adjusts the tone of the response accordingly.
[1430] Input: User text and voice input.
[1431] Output: Virtual CEO's response (text and audio).
[1432] Step 12: Collect interaction data and retrain
[1433] The server collects user interaction logs and stores them in a database as interaction data. The collected data is used to retrain the emotion engine and generative AI models.
[1434] Input: Dialogue log between the user and the virtual CEO.
[1435] Output: Collected dialogue logs, retrained AI model, and emotion engine.
[1436] (Application Example 2)
[1437] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1438] In modern factory operations, diverse data needs to be collected and analyzed in real time, but systems for handling this data efficiently and effectively are not yet widespread. Furthermore, there is a lack of means to provide appropriate instructions and advice in real time based on worker emotions and production status, which can lead to problems such as decreased production efficiency and increased worker stress.
[1439] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting information related to the CEO, means for preprocessing the collected information to remove noise and unify the data format, means for training a generative AI model using the preprocessed data to generate a virtual CEO emulator, means for interacting with the virtual CEO through a user interface, means for collecting interaction data and continuously retraining the AI model, means for collecting and monitoring data on each machine and worker in the factory in real time, and means for analyzing production status and worker emotions based on the collected data, and for the virtual CEO to provide appropriate instructions and advice. This makes it possible to improve the efficiency of factory operations and reduce worker stress.
[1440] "Information related to the CEO" refers to publicly available information, internal documents, past email data, blog posts, interviews, and social media posts concerning the company's chief executive officer.
[1441] "Preprocessing" refers to a series of processes to remove noise from collected data and standardize the data format.
[1442] "Noise reduction" refers to the process of removing unnecessary elements (such as spelling mistakes, special characters, and HTML tags) from collected data.
[1443] "Data format standardization" refers to the process of converting data in different formats into a consistent format.
[1444] A "generative AI model" refers to an artificial intelligence model that has been trained using pre-collected data and is capable of generating appropriate responses to specific tasks.
[1445] A "virtual CEO emulator" refers to a virtual character based on a generative AI model trained to mimic the speech patterns and decision-making styles of a CEO.
[1446] "User interface" refers to the interactive means by which a user interacts with a system, including browsers, displays, etc.
[1447] "Interaction data" refers to data that records the dialogue logs and responses between the user and the virtual CEO emulator.
[1448] "Continuously retraining an AI model" refers to the process of periodically retraining an AI model to improve its accuracy based on collected interaction data.
[1449] "Data on each machine and worker within the factory" refers to real-time data such as the operating status of factory production equipment, the behavior of workers, and their emotional states.
[1450] "Real-time data collection and monitoring" refers to the process of instantly acquiring current conditions and continuously monitoring that data.
[1451] "Production status" refers to information related to the progress of product production within the factory and the operating status of machinery.
[1452] "Worker emotions" refers to data related to emotions, such as workers' stress levels, satisfaction levels, and fatigue levels.
[1453] "Providing instructions and advice" refers to the process by which the virtual CEO emulator recommends specific actions and measures based on the data it has collected.
[1454] This invention relates to a system using a virtual CEO emulator aimed at improving the efficiency of factory operations and reducing worker stress. The system has the following configuration:
[1455] Hardware and software to be used
[1456] The server will use a high-performance server, natural language processing libraries (such as SpaCy and NLTK), deep learning frameworks (such as TensorFlow and PyTorch), sentiment recognition libraries (such as openSMILE), and web scraping tools (such as BeautifulSoup and Scrapy).
[1457] The devices include IoT sensors within the factory, wearable devices for workers, tablets, and displays.
[1458] Program processing flow
[1459] Data collection
[1460] The server collects data in real time from IoT sensors and wearable devices within the factory. It also periodically collects relevant industry news and technical literature using web scraping tools.
[1461] Specific example:
[1462] "Workers' smartwatches measure heart rate and stress levels and transmit the data to a cloud server in real time."
[1463] "IoT sensors monitor the operating status of each machine and transmit the data to a server."
[1464] Data preprocessing
[1465] The server removes noise from the collected data and standardizes the data format. Specifically, it corrects typos and removes unnecessary tags and special characters.
[1466] Model Learning
[1467] The server uses pre-processed data to train generative AI models and emotion engines. This allows a virtual CEO emulator to understand the emotional state and productivity of workers and provide appropriate instructions and advice.
[1468] Generating a virtual CEO emulator
[1469] The server generates a virtual CEO emulator using the generated generative AI model and emotion engine. This emulator can interact in real time with workers and managers via tablets or displays.
[1470] Example of a prompt:
[1471] "Could you tell me the current status of the production line?"
[1472] "Monitor the stress levels of the workers and suggest breaks if necessary."
[1473] User Interface and Interaction
[1474] Users can receive real-time feedback and advice on production status and individual issues through interaction with a virtual CEO emulator. The server has the capability to analyze user input and emotions and generate appropriate responses.
[1475] Specific example:
[1476] "Workers ask a virtual CEO questions about the project's progress and receive specific advice. The server analyzes the user's questions and emotions in real time to generate the most appropriate response."
[1477] Collection of interaction data and retraining
[1478] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement. The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue.
[1479] In this way, the present invention improves the efficiency of factory operations and reduces stress on workers.
[1480] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1481] Step 1:
[1482] The server collects data in real time from IoT sensors and wearable devices within the factory. Inputs include data such as the operating status of each machine, workers' heart rates, and stress levels. Based on this, the server centrally stores this data in a database, making it immediately accessible.
[1483] Step 2:
[1484] The server preprocesses the collected data. The input is the raw data collected in step 1. Specifically, it corrects typos and grammatical errors, removes HTML tags and unnecessary special characters, and standardizes the data into a clean text format. The output is the clean data saved in the standardized format.
[1485] Step 3:
[1486] The server applies natural language processing techniques to the pre-processed data to extract important keywords and sentences. The input is the clean text data from step 2. Specific operations include tokenization, part-of-speech tagging, and sentence analysis. The output is a list of the extracted keywords and sentences.
[1487] Step 4:
[1488] The server trains a generative AI model using preprocessed data and natural language processing results. The input is the data from steps 2 and 3. A deep learning framework (such as TensorFlow or PyTorch) is used, and backpropagation is performed to optimize the model parameters. The output is the trained generative AI model.
[1489] Step 5:
[1490] The server trains its emotion engine using an emotion recognition library (such as openSMILE). The input consists of worker voice and text data. Its specific actions include voice analysis and emotion labeling. The output is the trained emotion engine.
[1491] Step 6:
[1492] The server generates a virtual CEO emulator using a trained generative AI model and an emotion engine. The input is the trained model from steps 4 and 5. The virtual CEO emulator can interact with the user in real time via a tablet or display. The output is the virtual CEO emulator running on the user interface.
[1493] Step 7:
[1494] The user begins interacting with a virtual CEO emulator. Specifically, they input prompts such as, "Please tell me the current status of the production line," or "Please monitor the stress levels of the workers and suggest breaks if necessary." The input can be text or voice from the user. The server analyzes this and generates the most appropriate response. The output is appropriate advice or instructions provided to the user.
[1495] Step 8:
[1496] The server collects user interaction data and stores it in the interaction database. The input is the interaction log from step 7. The output is the stored interaction data.
[1497] Step 9:
[1498] The server retrains the generative AI model and emotion engine using the collected interaction data. The input is the interaction data from step 8. As a result of the retraining, the accuracy of the AI model and emotion engine improves, enabling more natural and effective dialogue. The output is the further improved generative AI model and emotion engine.
[1499] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1500] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1501] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1502] [Fourth Embodiment]
[1503] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1504] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1505] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1506] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1507] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1508] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1509] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1510] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1511] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1512] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1513] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1514] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1515] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1516] System Overview
[1517] This invention relates to a system that generates a virtual CEO emulator by collecting and preprocessing information related to the CEO and training a generative AI model, enabling interaction through a user interface. This system can learn the CEO's statements and decisions, imitate them in real time, and interact with the user via a two-dimensional interface.
[1518] Program Processing Overview
[1519] This system consists of the following main phases:
[1520] 1. Data Collection Phase
[1521] 2. Data preprocessing phase
[1522] 3. Model training phase
[1523] 4. CEO Emulator Generation Phase
[1524] 5. Interaction Phase
[1525] Program operation description
[1526] Data collection phase
[1527] The server collects publicly available information related to the CEO using web scraping techniques. This includes social media posts, blog articles, interviews, and publicly available emails.
[1528] Users can improve the accuracy of data collection by uploading internal documents and files to the system.
[1529] Data preprocessing phase
[1530] The server preprocesses the collected data, removing noise and standardizing the data format. Specifically, it uses a grammar checking tool to correct typos and removes unnecessary HTML tags and special characters.
[1531] The server uses natural language processing technology to analyze text data and extract important keywords and sentences.
[1532] Model Learning Phase
[1533] The server uses a generative AI model (e.g., GPT-3 or BERT) to train on preprocessed data. Initial setup is performed, and the model is instantiated.
[1534] The server splits the data into a training set and a test set and performs training. During this process, it optimizes the model parameters using backpropagation.
[1535] The server evaluates the model's performance using a test set and retrains it as needed.
[1536] CEO Emulator Generation Phase
[1537] The server generates a virtual CEO emulator as a 2D interface. This emulator is equipped with speech recognition and speech synthesis capabilities, enabling real-time interaction with the user.
[1538] The server deploys the generated interface as a user interface on the cloud, making it accessible to users.
[1539] Interaction Phase
[1540] Users interact with a virtual CEO through their device. For example, they log in to a web application using a browser and open a dialogue screen.
[1541] The server analyzes user input (text or voice) and generates an appropriate response in real time. The generated response is then delivered to the user as voice output through a speech synthesis engine.
[1542] The server continuously collects new dialogue data and retrains the model, thereby improving the system's accuracy and response quality.
[1543] Specific example
[1544] Data collection examples
[1545] The server collects the CEO's social media posts and interview articles from the past five years and stores them in a database.
[1546] Users can supplement the collected data by uploading personal emails and internal documents to the system.
[1547] Data preprocessing example
[1548] The server analyzes the collected data and corrects typos and grammatical errors. It also removes unnecessary tags and special characters, converting the data into clean text.
[1549] Model Learning Example
[1550] The server uses pre-processed data to train a generative AI model, learning the CEO's speaking patterns. For example, it incorporates frequently used phrases and decision-making tendencies into the model.
[1551] Emulator generation example
[1552] The server generates a 2D virtual CEO and implements speech recognition and synthesis capabilities. These emulators are deployed on the cloud as the user interface.
[1553] User interaction examples
[1554] Users ask questions about project progress to a virtual CEO and receive specific advice. The server analyzes the user's questions in real time and generates the most appropriate answers.
[1555] In summary, this invention provides a system that allows CEOs to leverage their knowledge and judgment for the future. This will lead to increased efficiency in corporate operations and improved accuracy in decision-making.
[1556] The following describes the processing flow.
[1557] Step 1: Data Collection
[1558] The server uses web scraping techniques to collect data related to the CEO from specified websites and social media platforms. This includes publicly available blog posts, interviews, and social media posts by the CEO.
[1559] Users can use the system interface to upload additional CEO-related data, such as internal documents and past email data.
[1560] Step 2: Save
[1561] The server stores the collected data in a database for centralized management. The stored data includes various formats such as text, images, and audio files.
[1562] Step 3: Noise Reduction and Data Cleansing
[1563] The server performs noise reduction processing on the collected data. This includes tasks such as correcting spelling errors and removing unnecessary tags and special characters.
[1564] The server standardizes data formats and cleans up text data. For example, it removes HTML tags and converts it to plain text.
[1565] Step 4: Data analysis using natural language processing
[1566] The server uses natural language processing techniques to analyze clean text data. It performs tokenization, part-of-speech tagging, and sentence analysis to extract important keywords and sentences.
[1567] Step 5: Create a dataset
[1568] The server splits the analyzed data into a training set and a test set. For example, 70% of the data might be used for the training set and 30% for the test set.
[1569] Step 6: Training the Generative AI Model
[1570] The server trains a generative AI model using a training set. It uses deep learning algorithms and performs backpropagation to optimize the model's parameters.
[1571] The server trains the model multiple times, specifying the number of epochs, to improve the model's accuracy.
[1572] Step 7: Model Evaluation
[1573] The server evaluates the performance of the trained model using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[1574] The server retrains the model as needed to further improve its performance.
[1575] Step 8: Generating the CEO Emulator
[1576] The server uses the final generative AI model to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time.
[1577] Step 9: Deploying the User Interface
[1578] The server deploys the generated virtual CEO emulator to the cloud. It provides a URL that users can access via a browser, and functions as the user interface.
[1579] Step 10: Start interacting with the user
[1580] Users initiate a conversation with a virtual CEO using a web interface. They can ask questions and seek advice using text or voice input.
[1581] The server analyzes user input in real time and generates an appropriate response. The generated response is then delivered to the user as speech output using a speech synthesis engine.
[1582] Step 11: Collect interaction data and retrain
[1583] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement.
[1584] The server uses the collected interaction data to retrain the model and improve the quality of the dialogue.
[1585] (Example 1)
[1586] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1587] In today's business environment, there is a growing need to quickly and accurately mimic the judgments and instructions of managers in order to improve the efficiency of corporate operations and the precision of decision-making. Furthermore, there is an increasing need for systems that can utilize the knowledge and experience of managers even when they are physically absent. However, conventional technologies do not adequately provide methods for mimicking managers' thought patterns and statements in real time.
[1588] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1589] In this invention, the server includes means for collecting publicly available information, means for preprocessing the collected information, performing grammar checks and unifying data formats, means for training a generative AI model using the preprocessed data to generate a virtual manager emulator, means for interacting with the virtual manager through a user interface, and means for collecting interaction data and continuously retraining the AI model. This enables real-time imitation of the manager's knowledge and judgment, improving the efficiency of business operations and the accuracy of decision-making.
[1590] "Public information" refers to information that is accessible to a large number of people, such as information posted on websites, social media posts, blog articles, interview articles, and publicly available emails.
[1591] "Preprocessing" is the process of correcting typographical errors in collected data, removing unnecessary HTML tags and special characters, and standardizing data formats.
[1592] A "generative AI model" is an artificial intelligence model that is trained using collected and pre-processed data, and includes, for example, natural language generation models such as GPT-3.
[1593] A "virtual executive emulator" is a virtual character that mimics the statements and decisions of executives in real time based on a generated AI model, and it has speech recognition and speech synthesis capabilities.
[1594] "User interface" refers to the screens and applications that allow a user to interact with a virtual manager emulator, and includes web browsers and dedicated applications.
[1595] "Interaction data" refers to the data from conversations between the user and a virtual manager emulator, and this data is used to retrain the AI model.
[1596] "Natural language processing technology" refers to techniques for analyzing text data and extracting important keywords and sentences, and includes libraries such as NLTK and SpaCy.
[1597] "Speech recognition" is a technology that converts speech data into text data, and Google Speech-to-Text is an example of this.
[1598] "Speech synthesis" is a technology that converts text data into speech data, and Amazon Polly is an example of this.
[1599] This invention is a system that generates a virtual CEO emulator by collecting and pre-processing information related to the CEO and training a generated AI model, enabling interaction through a user interface. This system imitates the judgments and statements of executives in real time based on publicly available information, thereby improving the efficiency of corporate operations and the accuracy of decision-making.
[1600] The system configuration is mainly described in the following phases:
[1601] 1. Data Collection Phase
[1602] 2. Data preprocessing phase
[1603] 3. Model training phase
[1604] 4. CEO Emulator Generation Phase
[1605] 5. Interaction Phase
[1606] Data collection phase
[1607] The server uses web scraping techniques (e.g., Beautiful Soup) to collect publicly available information such as social media posts, blog articles, interviews, and published emails. Specifically, it uses Python libraries to crawl websites containing certain keywords (e.g., names of business owners, company names).
[1608] Users can improve the quality and quantity of data collection by uploading internal documents and personal emails to the system. The uploaded files are stored in a database on the server.
[1609] Data preprocessing phase
[1610] The server corrects the collected data using a grammar checking tool (e.g., Grammarly API) to eliminate typos and grammatical errors.
[1611] The server removes unnecessary HTML tags and special characters from the collected data, converting it into clean text data.
[1612] The server analyzes text data using natural language processing technologies (e.g., NLTK, SpaCy) to extract important keywords and sentences. For example, it can extract key topics and keywords from interview articles with business executives.
[1613] Model Learning Phase
[1614] The server instantiates a generative AI model (e.g., GPT-3) and loads pre-processed data. Instances of the generative AI model are created using the OpenAI API.
[1615] The server splits the preprocessed data into a training set and a test set, and then trains the model.
[1616] The server uses backpropagation to optimize the parameters of the generated AI model and retrains it to improve its accuracy.
[1617] CEO Emulator Generation Phase
[1618] The server generates a virtual manager emulator using a pre-trained generative AI model. This emulator is designed as a two-dimensional interface and includes speech recognition (e.g., Google Speech-to-Text) and speech synthesis (e.g., Amazon Polly) capabilities.
[1619] The server deploys a virtual manager emulator on the cloud and makes it accessible to users. For example, it might be deployed on AWS and configured to be accessible to users via a browser.
[1620] Interaction Phase
[1621] The user logs into the web application from their device and opens the interactive interface.
[1622] Example: A user logs into the system using a browser and asks a virtual manager, "What is the current project status?"
[1623] The server analyzes user input (text or voice) and generates an appropriate response.
[1624] Example: The server analyzes the user's question and provides the generated response as speech output using a speech synthesis engine.
[1625] The server continuously collects new interaction data and retrains the generative AI model. This improves the system's accuracy and response quality.
[1626] Examples of specific prompt statements include the following:
[1627] "Could you tell me the CEO's view on the progress of the current project?"
[1628] "We'd like to hear the CEO's opinion on our new marketing strategy."
[1629] As described above, the present invention makes it possible to mimic the knowledge and judgment of managers in real time, thereby improving the efficiency of corporate operations and the accuracy of decision-making.
[1630] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1631] The specific processing flow of the program
[1632] Step 1: Data Collection
[1633] The server launches a web scraping script to collect publicly available information such as social media posts, blog articles, interviews, and published emails. Specifically, it uses the Python Beautiful Soup library to crawl websites containing specific keywords.
[1634] Input: Website URL or specific keywords
[1635] Output: Collected text data (e.g., text from social media posts or blog articles)
[1636] Specific operation: The server accesses a specific website, retrieves the page content, and extracts the necessary parts.
[1637] Step 2: Data Upload
[1638] Users upload internal documents and personal emails to the system. The uploaded files are stored in a database on the server.
[1639] Input: A file selected by the user (e.g., PDF, text file)
[1640] Output: Text data stored in the database
[1641] Specific operation: The user sends a file to the server using the system's upload form, and the server saves its contents to the database.
[1642] Step 3: Data Cleaning
[1643] The server uses a grammar checking tool (e.g., Grammarly API) to correct typos and grammatical errors in the collected data.
[1644] Input: Collected text data
[1645] Output: Clean text data with typos and grammatical errors corrected.
[1646] Specific operation: The server sends text data to a grammar checking tool and receives the corrected text.
[1647] Step 4: Unify text formatting
[1648] The server removes unnecessary HTML tags and special characters from the collected data to standardize the data format.
[1649] Input: Text data with typos and grammatical errors corrected.
[1650] Output: Clean text data
[1651] Specific operation: The server uses regular expressions to remove HTML tags and special characters, converting the data into clean text.
[1652] Step 5: Natural Language Processing (NLP)
[1653] The server uses natural language processing techniques (e.g., NLTK, SpaCy) to analyze text data and extract important keywords and sentences.
[1654] Input: Clean text data
[1655] Output: Extracted keywords and important sentences
[1656] Specific operation: The server starts a natural language processing library, analyzes the text data, and extracts important information.
[1657] Step 6: Model instantiation and data loading
[1658] The server instantiates a generative AI model (e.g., GPT-3) and loads pre-processed data.
[1659] Input: Extracted keywords and important sentences
[1660] Output: Text data loaded as training data
[1661] Specific operation: The server uses the OpenAI API to create a GPT-3 instance and load the pre-processed data.
[1662] Step 7: Model Training
[1663] The server divides the pre-processed data into a training set and a test set, and then trains the generative AI model.
[1664] Input: Loaded training data
[1665] Output: Trained generative AI model
[1666] Specific operation: The server optimizes the model using training data and evaluates its accuracy using test data.
[1667] Step 8: Generate CEO emulator
[1668] The server generates a virtual manager emulator using a pre-trained generative AI model. This emulator has a two-dimensional interface and features speech recognition and speech synthesis capabilities.
[1669] Input: Trained generative AI model
[1670] Output: Virtual manager emulator
[1671] Specific operation: The server generates characters from the model and integrates speech recognition and speech synthesis engines.
[1672] Step 9: Interface Deployment
[1673] The server deploys a virtual manager emulator on the cloud, making it accessible to users.
[1674] Input: Virtual manager emulator
[1675] Output: Deployed user interface
[1676] Specific operation: The server deploys the emulator to a cloud service such as AWS and configures it so that users can access it via a browser.
[1677] Step 10: User Interaction
[1678] The user logs into the web application from their device, opens a conversational interface, and interacts with a virtual manager.
[1679] Input: User voice or text input
[1680] Output: Response from the virtual manager
[1681] Specific operation: The user inputs a question or instruction, and the server generates a response and sends it back.
[1682] Step 11: Data Collection and Retraining
[1683] The server continuously collects conversational data between the user and the virtual manager, and uses it to retrain the generated AI model.
[1684] Input: User interaction data
[1685] Output: Retrained generative AI model
[1686] Specific operation: The server collects interaction data, periodically retrains the model, and improves the system's accuracy and response quality.
[1687] Through the processing steps described above, this system can mimic the knowledge and judgment of managers in real time, thereby improving the efficiency of business operations and the accuracy of decision-making.
[1688] (Application Example 1)
[1689] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1690] In recent years, there has been a growing demand for increased efficiency in corporate management and improved accuracy in decision-making. However, means of passing on leaders' knowledge and judgment have been limited, and providing real-time advice has been difficult. In particular, there is a need for improved quality of immediate information provision and advice to users in virtual environments. Furthermore, when using virtual reality devices, the lack of appropriate conversational interfaces to enhance the user experience is a challenge.
[1691] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1692] In this invention, the server includes means for collecting information, means for preprocessing the collected information to remove noise and unify the data format, means for training a generative AI model using the preprocessed data to generate a virtual mentor emulator, means for interacting with the virtual mentor through a user interface, means for collecting interaction data and continuously retraining the AI model, and means for providing the user with advice on product selection within a virtual environment through a virtual reality device. This enables increased efficiency in business management and improved accuracy in decision-making, and provides a real-time conversational interface that enhances the user experience in a virtual environment.
[1693] "Information" refers to a series of data related to the user, such as documents, data, audio, and images.
[1694] "Collection" refers to the entire process of gathering information.
[1695] "Preprocessing" refers to the process of removing noise from collected information and standardizing the data format.
[1696] A "generative AI model" refers to an algorithm or system that uses a large dataset for artificial intelligence to learn and perform a specified task.
[1697] A "virtual leader emulator" refers to a virtual agent created to mimic the speech patterns and behavioral patterns of a specific leader.
[1698] A "user interface" refers to the interface through which a user interacts with a system.
[1699] "Interaction data" refers to data related to the interactions and operations that take place between the user and the system.
[1700] "Retraining an AI model" refers to the process of retraining an existing AI model to improve its performance based on new data.
[1701] "Virtual reality devices" refer to hardware and software used to provide a virtual reality (VR) environment.
[1702] "Product selection advice" refers to information and suggestions provided by a supervisor emulator to help users choose a specific product.
[1703] "Real-time" refers to the process of generating an immediate response to user input.
[1704] System Overview
[1705] This invention realizes a system that generates a virtual instructor emulator and provides advice to users through a virtual environment or virtual reality device. Specifically, users can move freely within a virtual store and receive real-time advice from a virtual instructor regarding product selection.
[1706] Hardware and software used
[1707] Server: Performs data collection, preprocessing, training and retraining of AI models, and generation of virtual leader emulators.
[1708] Virtual reality devices (HMDs, etc.): Used by users to move around within a virtual store and interact with a virtual instructor emulator through an interface.
[1709] Generative AI Models: Using GPT-3 or similar high-performance AI models, the system generates statements and advice from a virtual leader emulator.
[1710] Natural language processing tools: Used for analyzing and preprocessing collected text data.
[1711] Processing flow and data calculations
[1712] 1. Information Gathering: The server uses web scraping techniques to collect publicly available information and internal documents. This includes social media posts, interview articles, and publicly available emails.
[1713] 2. Preprocessing: The collected information is preprocessed to remove noise and standardize the data format. Specifically, grammar checking tools are used to correct typographical errors and unnecessary HTML tags and special characters are removed.
[1714] 3. Training Generative AI Models: GPT-3 or similar generative AI models are trained using preprocessed data. During training, the data is divided into a training set and a test set, and the model is optimized.
[1715] 4. Generation of a virtual leader emulator: The server generates a virtual leader emulator using a trained AI model. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time.
[1716] 5. Collection and retraining of interaction data: Data generated when users interact with the instructor emulator using a virtual reality device is collected and used to continuously retrain the AI model.
[1717] Specific example
[1718] In a virtual store:
[1719] The user wears an HMD (Head-Mounted Display) and moves around within the virtual store.
[1720] The participant makes the gesture of picking up any product and is asked, "What are the features of this product?"
[1721] A virtual coach emulator provides real-time responses such as, "This product utilizes the latest running technology, is extremely lightweight, and durable. It is especially suitable for long-distance running."
[1722] Example of a prompt
[1723] User: "Could you please give me some advice about this new smartphone?"
[1724] Virtual Leader Emulator: "This smartphone features the latest processor, making it extremely fast and offering excellent battery life. It's especially recommended for gaming and video editing."
[1725] Thus, this invention allows users to receive real-time advice from a mentor emulator when effectively selecting products in a virtual environment. This aims to improve the efficiency of business operations and the accuracy of decision-making.
[1726] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1727] Step 1:
[1728] The server collects information. Specifically, it uses web scraping techniques to collect publicly available information from the web (such as social media posts, blog articles, interviews, and publicly available emails) and stores it in a database. The input is a URL or API endpoint on the web, and the output is the collected text data.
[1729] Step 2:
[1730] The server preprocesses the collected information. Specifically, it uses a grammar checker to correct typos and removes HTML tags and special characters. Next, it analyzes the text data using natural language processing techniques to extract important keywords and sentences. The input is the collected text data, and the output is text data in a clean and consistent format.
[1731] Step 3:
[1732] The server trains a generative AI model using preprocessed data. Specifically, it uses a model such as GPT-3, training the model by splitting the data into training and test sets. It optimizes the model parameters using backpropagation. The input is preprocessed text data, and the output is the trained AI model.
[1733] Step 4:
[1734] The server generates a virtual leader emulator using a trained AI model. This emulator has speech recognition and speech synthesis capabilities and is deployed on the cloud. The input is the trained AI model, and the output is the virtual leader emulator.
[1735] Step 5:
[1736] The user navigates a virtual environment through a virtual reality device (HMD) and selects products. They use an interface to interact with a virtual instructor emulator and ask questions about the products. Input is either voice or text input from the user, and output is a voice response from the virtual instructor emulator.
[1737] Step 6:
[1738] The server collects interaction data between the user and the virtual instructor emulator. Specifically, it collects data including dialogue history and usage patterns, and uses this data to retrain the AI model. The input is interaction data, and the output is the retrained AI model.
[1739] Step 7:
[1740] The server deploys a retrained AI model to the cloud, improving the accuracy and responsiveness of the virtual mentor emulator. This allows users to receive up-to-date information and optimal advice in real time. The input is the retrained AI model, and the output is the improved virtual mentor emulator.
[1741] By following these steps, a system will be completed that improves the efficiency of business operations and the accuracy of decision-making.
[1742] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1743] System Overview
[1744] This invention relates to a system that generates a virtual CEO emulator by collecting and preprocessing information related to the CEO, and training a generative AI model and emotion engine, enabling interaction through a user interface. This system can recognize the user's emotions and adjust its responses based on those emotions.
[1745] Program Processing Overview
[1746] This system consists of the following main phases:
[1747] 1. Data Collection Phase
[1748] 2. Data preprocessing phase
[1749] 3. Model training phase
[1750] 4. Emotional Engine Learning Phase
[1751] 5. CEO Emulator Generation Phase
[1752] 6. Interaction Phase
[1753] Program operation description
[1754] Data collection phase
[1755] The server uses web scraping techniques to collect data related to the CEO from specified websites and social media platforms. This includes publicly available blog posts, interviews, and social media posts by the CEO.
[1756] Users can use the system interface to upload additional CEO-related data, such as internal documents and past email data.
[1757] Data storage
[1758] The server stores the collected data in a database for centralized management. The stored data includes various formats such as text, images, and audio files.
[1759] Noise reduction and data cleansing
[1760] The server performs noise reduction processing on the collected data. This includes tasks such as correcting spelling errors and removing unnecessary tags and special characters.
[1761] The server standardizes data formats and cleans up text data. For example, it removes HTML tags and converts it to plain text.
[1762] Data analysis using natural language processing
[1763] The server uses natural language processing techniques to analyze clean text data. It performs tokenization, part-of-speech tagging, and sentence analysis to extract important keywords and sentences.
[1764] Creating a dataset
[1765] The server splits the analyzed data into a training set and a test set. For example, 70% of the data might be used for the training set and 30% for the test set.
[1766] Training of generative AI models
[1767] The server trains a generative AI model using a training set. It uses deep learning algorithms and performs backpropagation to optimize the model's parameters.
[1768] The server trains the model multiple times, specifying the number of epochs, to improve the model's accuracy.
[1769] Learning the Emotion Engine
[1770] The server trains the emotion engine using a training set. It analyzes user voice and text input data, assigns emotion labels, and trains the model.
[1771] The server will implement a function that uses an emotion engine to recognize emotions in real time from user input.
[1772] Model evaluation
[1773] The server evaluates the performance of the trained generative AI models and sentiment engines using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[1774] The server retrains the model as needed to further improve its performance.
[1775] Creating a CEO emulator
[1776] The server uses the final generative AI model and emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time.
[1777] User interface deployment
[1778] The server deploys the generated virtual CEO emulator to the cloud. It provides a URL that users can access via a browser, and functions as the user interface.
[1779] Start interaction with the user
[1780] Users initiate a conversation with a virtual CEO using a web interface. They can ask questions and seek advice using text or voice input.
[1781] The server analyzes user input in real time and generates an appropriate response. The generated response is then delivered to the user as speech output using a speech synthesis engine.
[1782] The server analyzes the user's emotions using an emotion engine and adjusts its response based on the recognized emotions. For example, if the user is angry, it will respond in a calm tone.
[1783] Collection of interaction data and retraining
[1784] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement.
[1785] The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue.
[1786] Specific example
[1787] Data collection examples
[1788] The server collects the CEO's social media posts and interview articles from the past five years and stores them in a database.
[1789] Users can supplement the collected data by uploading personal emails and internal documents to the system.
[1790] Data preprocessing example
[1791] The server analyzes the collected data and corrects typos and grammatical errors. It also removes unnecessary tags and special characters, converting the data into clean text.
[1792] Model Learning Example
[1793] The server uses pre-processed data to train generative AI models and emotion engines, improving their ability to understand the CEO's speaking patterns and users' emotions.
[1794] Emulator generation example
[1795] The server generates a 2D virtual CEO and implements speech recognition and synthesis capabilities. These emulators are deployed on the cloud as the user interface.
[1796] User interaction examples
[1797] Users ask questions about project progress to a virtual CEO and receive specific advice. The server analyzes the user's questions and emotions in real time to generate the most appropriate answers.
[1798] Based on the above, the present invention provides a system that allows CEOs to leverage their knowledge and judgment into the future. This system recognizes user emotions and adjusts responses based on those emotions, enabling more natural and effective dialogue. This, in turn, improves the efficiency of corporate operations and the accuracy of decision-making.
[1799] The following describes the processing flow.
[1800] Step 1: Data Collection
[1801] The server collects data related to the CEO from specified websites and social media platforms using web scraping techniques. Specifically, it uses APIs to obtain social media posts and crawling techniques to obtain text data from blog posts and interviews.
[1802] Users upload additional CEO-related data, such as internal company documents and past email data, using the system interface. For example, they can upload files using a drag-and-drop function.
[1803] Step 2: Save
[1804] The server stores the collected data in a database for centralized management. For example, it might use a NoSQL database to store data in different formats (text, images, audio files, etc.).
[1805] Step 3: Noise Reduction and Data Cleansing
[1806] The server performs noise reduction on the collected data. This is done using automated scripts and natural language processing tools to correct spelling errors, remove unnecessary tags and special characters, and so on.
[1807] The server standardizes data formats and cleans up text data. For example, it removes HTML tags and converts it to plain text.
[1808] Step 4: Data analysis using natural language processing
[1809] The server uses natural language processing techniques to analyze clean text data. Specifically, it performs processes such as tokenization, part-of-speech tagging, sentence analysis, and semantic analysis.
[1810] The server extracts important keywords and sentences from the analysis results and stores them in a database. For example, it may use text frequency analysis or TF-IDF scoring.
[1811] Step 5: Create a dataset
[1812] The server splits the analyzed data into a training set and a test set. For example, 70% of the data could be used for the training set and 30% for the test set. The script is then executed to distribute the data randomly.
[1813] Step 6: Training the Generative AI Model
[1814] The server trains a generative AI model (e.g., GPT-3 or BERT) using the training set. It then performs backpropagation to optimize the model's parameters using a deep learning algorithm.
[1815] The server trains the model multiple times, specifying the number of epochs (the number of training iterations), to improve the model's accuracy. Batch processing is used to perform calculations efficiently during this process.
[1816] Step 7: Learning the Emotion Engine
[1817] The server simultaneously trains the emotion engine using the training set. It analyzes user voice and text input data, assigns emotion labels, and trains the model. Specifically, it performs supervised learning using labeled datasets.
[1818] The server implements a function to recognize emotions in real time from user input using an emotion engine. For example, it uses speech feature extraction for emotion analysis from speech and emotion classification algorithms for emotion analysis from text.
[1819] Step 8: Model Evaluation
[1820] The server evaluates the performance of the trained generative AI models and sentiment engines using a test set. It calculates evaluation metrics such as accuracy, recall, and F1 score to verify the model's performance.
[1821] The server retrains the model as needed to further improve its performance, for example, by performing hyperparameter tuning.
[1822] Step 9: Generate the CEO emulator
[1823] The server uses the final generative AI model and emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition and speech synthesis capabilities and can interact with the user in real time. A graphical user interface (GUI) is designed to be user-friendly.
[1824] Step 10: Deploying the User Interface
[1825] The server deploys the generated virtual CEO emulator to the cloud. For example, it uses a web application framework to provide a URL that users can access through a browser, thus functioning as a user interface.
[1826] Step 11: Start interacting with the user
[1827] Users initiate a conversation with a virtual CEO using a web interface. They can ask questions and seek advice using text or voice input. The service operates in a browser and uses a microphone for voice input.
[1828] The server analyzes user input (text or voice) in real time and generates an appropriate response. The generated response is then delivered to the user as voice output using a speech synthesis engine. For example, natural language generation (NLG) technology is used to generate text, and a text-to-speech (TTS) engine is used to generate speech.
[1829] The server analyzes the user's emotions using an emotion engine and adjusts its response based on the recognized emotions. For example, if the user is angry, it will respond in a calm tone; if the user is sad, it will respond in a comforting tone.
[1830] Step 12: Collect interaction data and retrain
[1831] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement. The interaction logs stored in the database include user questions, virtual CEO responses, and user sentiment information.
[1832] The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue. Regular retraining allows the system to evolve over time, enabling more accurate responses.
[1833] (Example 2)
[1834] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1835] In current corporate operations, there is a lack of means to virtually replicate the knowledge and judgment of the CEO and to facilitate effective communication with employees and stakeholders. Furthermore, there is no system that can recognize user emotions in real time and adjust responses accordingly. This can potentially lead to a decrease in the efficiency of decision-making and problem-solving within companies.
[1836] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1837] In this invention, the server includes means for collecting information related to the CEO, means for preprocessing the collected information to remove noise and standardize the data format, means for analyzing the preprocessed data using natural language processing techniques to extract important keywords and sentences, means for training a generative AI model using the extracted data, means for generating a virtual CEO emulator with emotion recognition capabilities using the generated AI model, means for interacting with the virtual CEO through a user interface, means for collecting interaction data and continuously retraining the AI model, means for recognizing emotions in real time from user input, and means for adjusting responses based on recognized emotions. This makes it possible to virtually reproduce the knowledge and judgment of the CEO and generate appropriate responses according to the user's emotions.
[1838] "CEO-related information" refers to text, image, and audio data related to the CEO, such as blog posts, interviews, social media posts, internal documents, and past email data.
[1839] "Means of collection" refers to systems that include methods for obtaining specified information from the internet using web scraping technology, as well as methods for uploading internal documents and past email data via a user interface.
[1840] "Methods for preprocessing, noise reduction, and data format standardization" refer to techniques for processing collected data to create a clean data format by correcting spelling errors, removing unnecessary tags and special characters, and converting HTML tags to plain text.
[1841] "Natural language processing technology" refers to algorithms and methods used to analyze text data and extract important keywords and sentences through tokenization, part-of-speech tagging, and sentence analysis.
[1842] A "generative AI model" is an AI model trained using deep learning algorithms that has the ability to generate appropriate outputs for specific inputs.
[1843] A "virtual CEO emulator" is a virtual character created based on a generative AI model, possessing speech recognition and speech synthesis capabilities, and serving as an interface that can interact with users in real time.
[1844] A "user interface" refers to the graphical screen or web page that a user uses to interact with a virtual CEO.
[1845] "Interaction data" refers to digital information, including dialogue logs and response data from interactions between the user and the virtual CEO emulator.
[1846] "Emotion recognition functionality" refers to technology that analyzes user voice and text input data and identifies the user's emotions based on emotion labels extracted from that data.
[1847] "Means of adjusting responses based on emotions" refers to methods or algorithms for adjusting the content and tone of generated responses according to the user's emotions identified by emotion recognition functions.
[1848] This invention is a system that collects information related to a CEO, trains a generative AI model based on that information to generate a virtual CEO emulator, and allows the user to interact with it through an interface. This system also has the ability to recognize the user's emotions and adjust its responses based on those emotions. The specific embodiments of this system are described in detail below.
[1849] Data collection and storage
[1850] The server collects information related to the CEO from specified websites and social media platforms. This task utilizes web scraping techniques such as BeautifulSoup and Scrapy. The collected information includes the CEO's blog posts, interviews, and social media posts. The collected data is stored in databases such as MySQL and MongoDB.
[1851] Specific example:
[1852] The server collects social media posts and interview articles from the past five years and stores them in a database.
[1853] Users upload internal documents and past email data using the system interface.
[1854] Data preprocessing
[1855] The server performs noise reduction and data formatting on the collected data. Specifically, it corrects spelling errors using regular expressions and removes unnecessary HTML tags and special characters. Text data is converted to plain text. Python is often used for this process.
[1856] Analysis using natural language processing
[1857] The server applies natural language processing techniques to the pre-processed data. For example, NLTK or SpaCy are used to tokenize text data, tag parts of speech, and analyze sentences, extracting important keywords and sentences.
[1858] Specific example:
[1859] The server tokenizes the text data, tags it with parts of speech, and extracts important keywords.
[1860] Model training
[1861] The server trains generative AI models using the extracted data. Deep learning frameworks used include TensorFlow and PyTorch. Parameter optimization is performed using backpropagation with the training set.
[1862] Learning the Emotion Engine
[1863] The server trains an emotion engine, as well as a trained generative AI model. This emotion engine uses a BERT-based emotion analysis model to extract emotion labels (e.g., joy, sadness, anger) from the user's voice and text input data.
[1864] Specific example:
[1865] The server analyzes the audio data and trains a model to assign emotion labels.
[1866] Creating a virtual emulator
[1867] The server integrates a generative AI model and an emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition (Google Speech-to-Text API) and speech synthesis (Amazon Polly) capabilities.
[1868] Deployment and User Interaction
[1869] The server deploys the generated virtual CEO emulator to a cloud environment (AWS or Azure), making it accessible to users via a browser. Users can initiate an interaction with the virtual CEO using the interface. The server analyzes questions entered via text or voice in real time and generates appropriate responses. These responses are output as speech using a speech synthesis engine. It also features an emotion engine that analyzes the user's emotions and adjusts the tone of the responses accordingly.
[1870] Specific example:
[1871] The user enters, "Based on the CEO's latest social media posts, please share your insights into current market trends."
[1872] The server analyzes this input and generates a response through a virtual CEO emulator.
[1873] Collection of interaction data and retraining
[1874] The server collects user interaction logs and stores them as interaction data. Based on this data, the model is retrained to continuously improve performance.
[1875] The above describes a specific embodiment for implementing the present invention, which will lead to increased efficiency in decision-making and problem-solving within companies.
[1876] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1877] Step 1: Data Collection
[1878] The server collects information related to the CEO from specified websites and social media platforms. Specifically, it uses web scraping techniques (e.g., BeautifulSoup, Scrapy) to obtain the data.
[1879] Users upload internal documents and past email data through the system interface.
[1880] Input: Website URL, social media platform account information, user-uploaded internal documents and email data.
[1881] Output: Collected text data, image data, and audio data.
[1882] Step 2: Save Data
[1883] The server stores the collected data in a database (e.g., MySQL, MongoDB). Text data, image data, audio data, etc., are stored in the appropriate format according to their respective requirements.
[1884] Input: Collected text data, image data, and audio data.
[1885] Output: Clean data stored in the database.
[1886] Step 3: Noise Reduction and Data Cleansing
[1887] The server performs noise reduction processing on the collected data. Specifically, it uses regular expressions to correct spelling mistakes and remove unnecessary HTML tags and special characters.
[1888] Input: Raw data stored in the database.
[1889] Output: Denoised and cleaned text data.
[1890] Step 4: Data analysis using natural language processing
[1891] The server uses pre-processed data to apply natural language processing techniques (e.g., NLTK, SpaCy) to extract important keywords and sentences through tokenization, part-of-speech tagging, and sentence analysis.
[1892] Input: Clean text data.
[1893] Output: Tokenized keywords, part-of-speech tags, and parsed sentences.
[1894] Step 5: Create a dataset
[1895] The server splits the analyzed data into a training set (e.g., 70% of the data) and a test set (e.g., 30% of the data). These datasets are used for model training and evaluation.
[1896] Input: Analyzed keywords and sentences.
[1897] Output: Training set, test set.
[1898] Step 6: Training the Generative AI Model
[1899] The server trains a generative AI model (e.g., GPT-3) using a training set. Deep learning frameworks used include TensorFlow and PyTorch. Parameter optimization is performed using backpropagation.
[1900] Input: Training set.
[1901] Output: A trained generative AI model.
[1902] Step 7: Learning the Emotion Engine
[1903] The server trains an emotion engine (e.g., a BERT-based emotion analysis model) using a training set. The model is trained by assigning emotion labels to user voice and text input data.
[1904] Input: Training set, user voice and text data.
[1905] Output: Trained emotion engine.
[1906] Step 8: Model Evaluation
[1907] The server evaluates the performance of generative AI models and emotion engines trained using a test set. Evaluation metrics include accuracy, recall, and F1 score, and retraining is performed as needed.
[1908] Input: Test set.
[1909] Output: Evaluation results (accuracy, recall, F1 score).
[1910] Step 9: Generate the CEO emulator
[1911] The server integrates the final generative AI model and emotion engine to generate a virtual CEO emulator with a 2D interface. This emulator has speech recognition (Google Speech-to-Text API) and speech synthesis (Amazon Polly) capabilities.
[1912] Input: Trained generative AI model, emotion engine.
[1913] Output: Virtual CEO emulator.
[1914] Step 10: Deploying the User Interface
[1915] The server deploys the generated virtual CEO emulator to a cloud environment (AWS or Azure) and makes it accessible to users through a browser.
[1916] Input: Virtual CEO emulator.
[1917] Output: Deployed interface (URL).
[1918] Step 11: Interacting with the user
[1919] Users initiate a conversation with a virtual CEO through their browser, entering questions and concerns via text or voice.
[1920] The server analyzes user input in real time and generates an appropriate response. The generated response is delivered to the user using a speech synthesis engine. Additionally, an emotion engine analyzes the user's emotions and adjusts the tone of the response accordingly.
[1921] Input: User text and voice input.
[1922] Output: Virtual CEO's response (text and audio).
[1923] Step 12: Collect interaction data and retrain
[1924] The server collects user interaction logs and stores them in a database as interaction data. The collected data is used to retrain the emotion engine and generative AI models.
[1925] Input: Dialogue log between the user and the virtual CEO.
[1926] Output: Collected dialogue logs, retrained AI model, and emotion engine.
[1927] (Application Example 2)
[1928] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1929] In modern factory operations, diverse data needs to be collected and analyzed in real time, but systems for handling this data efficiently and effectively are not yet widespread. Furthermore, there is a lack of means to provide appropriate instructions and advice in real time based on worker emotions and production status, which can lead to problems such as decreased production efficiency and increased worker stress.
[1930] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting information related to the CEO, means for preprocessing the collected information to remove noise and unify the data format, means for training a generative AI model using the preprocessed data to generate a virtual CEO emulator, means for interacting with the virtual CEO through a user interface, means for collecting interaction data and continuously retraining the AI model, means for collecting and monitoring data on each machine and worker in the factory in real time, and means for analyzing production status and worker emotions based on the collected data, and for the virtual CEO to provide appropriate instructions and advice. This makes it possible to improve the efficiency of factory operations and reduce worker stress.
[1931] "Information related to the CEO" refers to publicly available information, internal documents, past email data, blog posts, interviews, and social media posts concerning the company's chief executive officer.
[1932] "Preprocessing" refers to a series of processes to remove noise from collected data and standardize the data format.
[1933] "Noise reduction" refers to the process of removing unnecessary elements (such as spelling mistakes, special characters, and HTML tags) from collected data.
[1934] "Data format standardization" refers to the process of converting data in different formats into a consistent format.
[1935] A "generative AI model" refers to an artificial intelligence model that has been trained using pre-collected data and is capable of generating appropriate responses to specific tasks.
[1936] A "virtual CEO emulator" refers to a virtual character based on a generative AI model trained to mimic the speech patterns and decision-making styles of a CEO.
[1937] "User interface" refers to the interactive means by which a user interacts with a system, including browsers, displays, etc.
[1938] "Interaction data" refers to data that records the dialogue logs and responses between the user and the virtual CEO emulator.
[1939] "Continuously retraining an AI model" refers to the process of periodically retraining an AI model to improve its accuracy based on collected interaction data.
[1940] "Data on each machine and worker within the factory" refers to real-time data such as the operating status of factory production equipment, the behavior of workers, and their emotional states.
[1941] "Real-time data collection and monitoring" refers to the process of instantly acquiring current conditions and continuously monitoring that data.
[1942] "Production status" refers to information related to the progress of product production within the factory and the operating status of machinery.
[1943] "Worker emotions" refers to data related to emotions, such as workers' stress levels, satisfaction levels, and fatigue levels.
[1944] "Providing instructions and advice" refers to the process by which the virtual CEO emulator recommends specific actions and measures based on the data it has collected.
[1945] This invention relates to a system using a virtual CEO emulator aimed at improving the efficiency of factory operations and reducing worker stress. The system has the following configuration:
[1946] Hardware and software to be used
[1947] The server will use a high-performance server, natural language processing libraries (such as SpaCy and NLTK), deep learning frameworks (such as TensorFlow and PyTorch), sentiment recognition libraries (such as openSMILE), and web scraping tools (such as BeautifulSoup and Scrapy).
[1948] The devices include IoT sensors within the factory, wearable devices for workers, tablets, and displays.
[1949] Program processing flow
[1950] Data collection
[1951] The server collects data in real time from IoT sensors and wearable devices within the factory. It also periodically collects relevant industry news and technical literature using web scraping tools.
[1952] Specific example:
[1953] "Workers' smartwatches measure heart rate and stress levels and transmit the data to a cloud server in real time."
[1954] "IoT sensors monitor the operating status of each machine and transmit the data to a server."
[1955] Data preprocessing
[1956] The server removes noise from the collected data and standardizes the data format. Specifically, it corrects typos and removes unnecessary tags and special characters.
[1957] Model Learning
[1958] The server uses pre-processed data to train generative AI models and emotion engines. This allows a virtual CEO emulator to understand the emotional state and productivity of workers and provide appropriate instructions and advice.
[1959] Generating a virtual CEO emulator
[1960] The server generates a virtual CEO emulator using the generated generative AI model and emotion engine. This emulator can interact in real time with workers and managers via tablets or displays.
[1961] Example of a prompt:
[1962] "Could you tell me the current status of the production line?"
[1963] "Monitor the stress levels of the workers and suggest breaks if necessary."
[1964] User Interface and Interaction
[1965] Users can receive real-time feedback and advice on production status and individual issues through interaction with a virtual CEO emulator. The server has the capability to analyze user input and emotions and generate appropriate responses.
[1966] Specific example:
[1967] "Workers ask a virtual CEO questions about the project's progress and receive specific advice. The server analyzes the user's questions and emotions in real time to generate the most appropriate response."
[1968] Collection of interaction data and retraining
[1969] The server collects user interaction logs and stores them as interaction data. This enables continuous system improvement and accuracy enhancement. The server uses the collected interaction data to retrain the emotion engine and generative AI models, improving the quality of the dialogue.
[1970] In this way, the present invention improves the efficiency of factory operations and reduces stress on workers.
[1971] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1972] Step 1:
[1973] The server collects data in real time from IoT sensors and wearable devices within the factory. Inputs include data such as the operating status of each machine, workers' heart rates, and stress levels. Based on this, the server centrally stores this data in a database, making it immediately accessible.
[1974] Step 2:
[1975] The server preprocesses the collected data. The input is the raw data collected in step 1. Specifically, it corrects typos and grammatical errors, removes HTML tags and unnecessary special characters, and standardizes the data into a clean text format. The output is the clean data saved in the standardized format.
[1976] Step 3:
[1977] The server applies natural language processing techniques to the pre-processed data to extract important keywords and sentences. The input is the clean text data from step 2. Specific operations include tokenization, part-of-speech tagging, and sentence analysis. The output is a list of the extracted keywords and sentences.
[1978] Step 4:
[1979] The server trains a generative AI model using preprocessed data and natural language processing results. The input is the data from steps 2 and 3. A deep learning framework (such as TensorFlow or PyTorch) is used, and backpropagation is performed to optimize the model parameters. The output is the trained generative AI model.
[1980] Step 5:
[1981] The server trains its emotion engine using an emotion recognition library (such as openSMILE). The input consists of worker voice and text data. Its specific actions include voice analysis and emotion labeling. The output is the trained emotion engine.
[1982] Step 6:
[1983] The server generates a virtual CEO emulator using a trained generative AI model and an emotion engine. The input is the trained model from steps 4 and 5. The virtual CEO emulator can interact with the user in real time via a tablet or display. The output is the virtual CEO emulator running on the user interface.
[1984] Step 7:
[1985] The user begins interacting with a virtual CEO emulator. Specifically, they input prompts such as, "Please tell me the current status of the production line," or "Please monitor the stress levels of the workers and suggest breaks if necessary." The input can be text or voice from the user. The server analyzes this and generates the most appropriate response. The output is appropriate advice or instructions provided to the user.
[1986] Step 8:
[1987] The server collects user interaction data and stores it in the interaction database. The input is the interaction log from step 7. The output is the stored interaction data.
[1988] Step 9:
[1989] The server retrains the generative AI model and emotion engine using the collected interaction data. The input is the interaction data from step 8. As a result of the retraining, the accuracy of the AI model and emotion engine improves, enabling more natural and effective dialogue. The output is the further improved generative AI model and emotion engine.
[1990] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1991] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1992] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1993] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1994] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1995] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1996] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1997] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1998] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1999] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emo...
Claims
1. Means of collecting information related to the CEO, A means for preprocessing the collected information, removing noise, and standardizing the data format, A means of training a generative AI model using preprocessed data to generate a virtual CEO emulator, A means of interacting with a virtual CEO through a user interface, A system that includes means for collecting interaction data and continuously retraining AI models.
2. The system according to claim 1, comprising means for analyzing preprocessed data using natural language processing techniques and extracting important keywords and sentences.
3. The system according to claim 1, comprising means for generating a virtual CEO emulator as a two-dimensional interface and interacting with a user through speech recognition and speech synthesis functions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A