System
The system addresses the challenge of integrating and evaluating public information by collecting, preprocessing, and training a generative AI model to provide unified and reliable data, enhancing the efficiency and accuracy of information provision.
Patent Information
- Application Number
- JP2024118213
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
The vast amount of public information available is not effectively utilized, and there is a lack of systems to centrally integrate and evaluate its accuracy and reliability, particularly in policy formulation and life support, leading to inefficiencies and difficulties in obtaining reliable information.
A system that collects data from public sources, preprocesses it, stores it in a database, trains a generative AI model, maps multiple perspectives into a unified space, receives user requests, and generates and provides integrated information using a generative AI model.
Enables centralized management and efficient provision of highly reliable information, allowing users to quickly access the latest and most accurate data from diverse sources.
Smart Images

Figure 2026017431000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, the amount of public information available is constantly increasing, but there is a lack of effective ways to utilize this vast amount of information. Furthermore, information obtained from diverse sources is fragmented, and there is no way to centrally integrate it, making it difficult to evaluate the accuracy and reliability of the information. This issue is particularly pronounced in policy formulation and life support, and requires a great deal of time and effort. Therefore, there is a need for an innovative system that can efficiently and accurately integrate public information and analyze and propose solutions from various perspectives. [Means for solving the problem]
[0005] To solve the above problems, the present invention provides the following means. A system is configured including: means for collecting data from public information sources; means for preprocessing the collected data; means for storing the preprocessed data in a database; means for inputting the stored data into a generative AI model; means for training the generative AI model; means for mapping multiple perspectives into a unified space from the trained generative AI model; means for receiving requests from users; means for searching and generating information based on the requests; and means for providing the generated information to users. This system enables centralized management of diverse data collected from information sources and the provision of information that integrates multiple perspectives. Furthermore, the use of a generative AI model enables the generation and provision of highly reliable information.
[0006] "Public sources" are sources of information such as government databases, news sites, and social media that provide data and information that is widely available on the internet.
[0007] "Data collection methods" refers to the technologies and methods used to obtain the required data from public sources, including, for example, the use of APIs and web scraping.
[0008] "Preprocessing means" is a general term for techniques and methods for analyzing and formatting collected data, including data cleaning and format standardization.
[0009] "Means of storing data in a database" refers to systems and technologies for efficiently and securely storing pre-processed data, including database systems and storage solutions.
[0010] A "generative AI model" is an artificial intelligence model trained using machine learning algorithms, and has the ability to make inferences and predictions on new data.
[0011] "Training means" refers to the techniques and methods used to train a generative AI model using collected data, including running machine learning algorithms.
[0012] "Means of mapping multiple perspectives into a unified space" is a general term for techniques and methods for centrally aggregating data obtained from different information sources and perspectives and visualizing the relative positions and connections between the perspectives.
[0013] "Means for receiving user requests" refers to the technology or system for receiving information requests from users through a user interface or API.
[0014] "Means for searching and generating information" is a general term for technologies and methods for searching for relevant information in a database based on a user request and generating new information.
[0015] "Means of providing information to users" refers to the technologies or systems used to display or transmit generated information to users, including web interfaces and mobile apps. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] The present invention is a system that trains a generative AI model using data collected from public information sources and provides information from a unified perspective. This system consists of a server, a terminal, and a user.
[0038] Public information collection and preprocessing
[0039] First, the server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, it obtains the necessary data using APIs and web scraping technology. For example, it collects the latest policy information on COVID-19 from the government's open data portal and related articles from news sites. Next, the server preprocesses the collected data. It cleans the data, standardizes the format, removes special characters and spaces, and arranges the data into a unified format.
[0040] Storing in a database and training the generated AI model
[0041] The preprocessed data is stored in a database by the server. The server then retrieves the latest data from the database and inputs it into the generative AI model. This generative AI model has the ability to learn from the collected data and make inferences and predictions about new data. The server trains the generative AI model so that it always reflects the latest information.
[0042] Mapping to a unified space and calculating the relative viewpoints
[0043] The server runs a technology that maps multiple viewpoints into a unified space using a trained generative AI model. During this process, data from different sources and viewpoints is centrally aggregated. Furthermore, a clustering algorithm is used to calculate and visualize the relative positions of each viewpoint. As a result, users can easily grasp the relevance and reliability of the information.
[0044] Receiving user requests and generating information
[0045] A user requests specific information through a device (such as a PC or smartphone). For example, they send a request such as, "Please tell me the latest measures against COVID-19." The device then sends this request to the server. The server receives the request from the user and analyzes it using natural language processing (NLP). Based on the analyzed request, the server searches for relevant information from a unified space and generates appropriate comments or policy proposals using a generative AI model.
[0046] Providing information and concrete examples
[0047] The server sends the generated information to the device, which then displays the results to the user. For example, if a user requests, "I want to know about COVID-19 countermeasures," the server generates the latest policy information and suggestions, and sends information such as, "The latest government policy recommends vaccination in XX region. In addition, the home quarantine period has been changed from XX to XX days," to the device, which then displays this to the user. This system allows users to always easily access the latest, most reliable information, which can be used to understand policies and support their daily lives.
[0048] The processing flow will be explained below.
[0049] Step 1:
[0050] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.) by using REST APIs to retrieve data in JSON format or by using web scraping technology to obtain the required information.
[0051] Step 2:
[0052] The server pre-processes the data it receives, cleaning it (for example, removing whitespace and special characters), standardizing date and number formats, and arranging the data into a consistent format.
[0053] Step 3:
[0054] The server saves the pre-processed data to a database, where the information is stored efficiently and securely using INSERT operations into an SQL database.
[0055] Step 4:
[0056] The server retrieves the latest data from the database and inputs it into the generative AI model, which then uses the newly collected data to train the generative AI model.
[0057] Step 5:
[0058] The server trains the generative AI model, running machine learning algorithms that enable the model to make accurate predictions and inferences on new data.
[0059] Step 6:
[0060] The server maps multiple perspectives from the generative AI model into a unified space. Specifically, it plots data from different sources and perspectives in a multidimensional space and calculates their spatial relationships using a clustering algorithm.
[0061] Step 7:
[0062] The user requests specific information through the device, for example, "Please tell me the latest information on COVID-19 countermeasures," and sends it.
[0063] Step 8:
[0064] The terminal sends a request from the user to the server, and HTTP requests are generally used as the communication protocol.
[0065] Step 9:
[0066] The server analyzes the request received from the terminal, and uses natural language processing (NLP) to analyze the content of the request and identify its requirements.
[0067] Step 10:
[0068] The server searches for relevant information from the unified space based on the analysis results, leveraging the capabilities of SQL queries and generative AI models to extract the necessary information.
[0069] Step 11:
[0070] The server uses generative AI models to generate comments and policy proposals based on the retrieved information, using natural language generation (NLG) to provide the generated text in a format that is easy for users to understand.
[0071] Step 12:
[0072] The server then sends the generated information to the terminal. Again, HTTP responses are generally used as the communication protocol.
[0073] Step 13:
[0074] The device displays the information received from the server to the user, either by updating a web page or displaying the information in the mobile app UI, making it easily accessible to the user.
[0075] This series of steps allows users to efficiently utilize open public information and obtain the information they need quickly and accurately.
[0076] Example 1
[0077] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0078] In conventional information provision systems, when collecting data from public information sources and providing information from a unified perspective, multiple processes such as data preprocessing, training of generative AI models, and viewpoint mapping are required, and each process is often performed independently, resulting in problems that reduce the efficiency and accuracy of the entire system.In addition, there is a lack of functionality to properly calculate and visualize the positional relationships of viewpoints obtained from different information sources, making it difficult for users to quickly obtain reliable information.
[0079] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0080] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a data management device, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple viewpoints into a unified space from the trained generative AI model, means for calculating the positional relationship between viewpoints obtained from different information sources, means for receiving a request from a user, means for searching and generating information based on the request, and means for providing the generated information to the user, thereby enabling the user to quickly obtain reliable information and grasp a variety of information in a unified manner.
[0081] A "public information source" is a database or website that provides information that is publicly accessible on the Internet.
[0082] "Data preprocessing" is the process of cleaning, formatting, and removing unnecessary information from collected data to prepare it in a format that is easy to analyze and store.
[0083] A "data management device" refers to a database or storage system for efficiently and safely storing collected data.
[0084] A "generative artificial intelligence model" is a machine learning model that can learn from collected data and predict or generate data in the future.
[0085] "Training" is the learning process that a generative artificial intelligence model goes through to achieve optimal performance based on data.
[0086] A "unified space" is a data representation that centrally manages multiple perspectives collected from different sources and maps them in a form that makes comparison and analysis easy.
[0087] "Positional relationships between viewpoints" refers to information that indicates how multiple viewpoints are related to each other.
[0088] A "clustering algorithm" is a computational method for grouping data based on similarity.
[0089] "User Request" means a request for information entered by a user into the system.
[0090] "Information retrieval and generation means" refers to the process of locating appropriate information from databases and generative artificial intelligence models based on user requests and generating text, reports, etc. as needed.
[0091] "Generated Information" refers to information created by the system using a generative artificial intelligence model based on a user request.
[0092] The present invention is a system that trains a generative AI model using data collected from public information sources and provides information from a unified perspective. This system consists of a server, a terminal, and a user.
[0093] First, the server collects data from public sources on the internet, such as government databases, news sites, and social networking services. This process utilizes APIs and web scraping techniques, using tools such as Python's requests library and BeautifulSoup. For example, it obtains policy information about COVID-19 from government open data portals and collects related articles from news sites.
[0094] The server then preprocesses the collected data, which includes removing special characters and whitespace, cleaning, and formatting, using the Python pandas library to organize the data into a unified format.
[0095] The preprocessed data is stored in a data management device by the server. The server uses a database system such as MySQL or PostgreSQL to efficiently store the data. The stored data is then input into a generative artificial intelligence model to train the model. Machine learning frameworks such as TensorFlow and PyTorch are used for this training.
[0096] A trained generative artificial intelligence model maps multiple viewpoints into a unified space. The server integrates data from different sources and viewpoints and calculates the spatial relationships between viewpoints using a clustering algorithm in Python's Scikit-learn library. As a result, the relationships between viewpoints are visualized, and important information is centralized.
[0097] A user requests specific information through a device, such as a PC or smartphone. An example of a prompt would be "Please tell me the latest measures against COVID-19." The device then sends this request to the server, which analyzes it using natural language processing (NLP) tools. The server uses Python's spaCy or NLTK to analyze the request, searches for relevant information in a unified space based on the analysis results, and generates comments or policy proposals using a generative artificial intelligence model if necessary.
[0098] Finally, the server sends the generated information to the device, which then displays the results to the user. For example, information such as "The latest government policy recommends vaccination in the ____ region. Furthermore, the self-isolation period has been changed from ____ to ____ days" can be provided. In this way, users can quickly obtain the latest and most reliable information.
[0099] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0100] Step 1: Gather public information
[0101] The server collects data from public sources on the Internet, including government databases, news sites, social networking services, etc. The server uses Python's requests library to retrieve the required data from APIs, and also uses BeautifulSoup to scrape articles from news sites.
[0102] Input: URL or API endpoint of a public source
[0103] Output: Raw data (text, JSON, etc.)
[0104] Step 2: Preprocessing the data
[0105] The server preprocesses the collected data, which includes removing unnecessary whitespace and special characters, standardizing the format, etc. The server uses the Python pandas library to clean the data and convert it into a unified format.
[0106] Input: Raw data
[0107] Output: Preprocessed data (unified format)
[0108] Step 3: Saving to the database
[0109] The server stores the preprocessed data in a data management device (database), such as MySQL or PostgreSQL. The server establishes a connection to the database and inserts data using SQL queries.
[0110] Input: Preprocessed data
[0111] Output: Data stored in the database
[0112] Step 4: Training the generative AI model
[0113] The server retrieves the latest data from the database and trains a generative AI model using machine learning frameworks such as TensorFlow and PyTorch. The server inputs the data into the model and trains it over multiple epochs.
[0114] Input: Data retrieved from the database
[0115] Output: A trained generative AI model
[0116] Step 5: Mapping to a unified space
[0117] The server uses a trained generative AI model to map multiple viewpoints into a unified space using a clustering algorithm using Python's Scikit-learn library, which calculates the spatial relationships between viewpoints obtained from different sources.
[0118] Input: A trained generative AI model and associated data
[0119] Output: viewpoint data mapped to a unified space
[0120] Step 6: Calculate the viewpoint position
[0121] The server calculates the relative positions of the viewpoints mapped to the unified space and prepares for visualization. It calculates the distance and relationship between viewpoints and generates a network graph. This process uses the Python networkx library.
[0122] Input: Mapped viewpoint data
[0123] Output: Data for visualizing the relative positions of viewpoints
[0124] Step 7: Receiving a user request
[0125] The user requests specific information through the device, for example by entering a prompt such as "Please tell me the latest information on COVID-19 countermeasures." The device then sends this request to the server.
[0126] Input: User request (prompt sentence)
[0127] Output: Request sent to the server
[0128] Step 8: Generate information
[0129] The server analyzes requests received from users using natural language processing (NLP) tools, such as Python's spaCy and NLTK. Based on the analysis results, the server searches for relevant information in a unified space and generates appropriate comments and policy proposals using a generative AI model.
[0130] Input: The request sent to the server
[0131] Output: Generated information (comments and policy suggestions)
[0132] Step 9: Provide information
[0133] The server sends the generated information to the device, which then displays the results to the user, such as, "The latest government policy recommends vaccination in the ____ region. In addition, the home quarantine period has been changed from ____ to ____ days."
[0134] Input: Generated information
[0135] Output: Information displayed on the user's terminal
[0136] In this way, specific actions are performed at each step, allowing the user to quickly obtain reliable, up-to-date information.
[0137] (Application example 1)
[0138] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0139] In today's world, businesses and individuals are constantly exposed to the threat of cyberattacks, resulting in increased risk of serious economic losses and information leaks. Cyberthreats evolve daily, and information about them is provided fragmentedly from numerous sources, making it difficult to unify, understand, and respond immediately. Furthermore, the lack of a system that can assess risks in real time and issue immediate alerts for high-risk threats means that prompt responses are delayed, potentially resulting in greater damage. To address these challenges and improve the effectiveness of cybersecurity, a system is needed that can unify the latest information, quickly and accurately assess risks, and notify users in real time.
[0140] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0141] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a database, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple perspectives into a unified space from the trained generative AI model, means for receiving requests from users, means for searching and generating information based on the requests, means for providing the generated information to users, means for risk assessing cybersecurity-related information, and means for issuing alerts to users about high-risk threats in real time. This enables rapid and accurate risk assessment of cyberattacks and the issuance of alerts in real time, thereby strengthening the security measures of companies and individuals.
[0142] "Public sources" refer to data sources that are accessible to anyone and from which information can be obtained, such as government databases, news sites, and social media.
[0143] "Data preprocessing" is the process of cleaning collected data, standardizing the format, removing unnecessary special characters, etc., to improve the quality of the data.
[0144] "Database" refers to a collection of data that stores collected and preprocessed data and structures it for easy access and retrieval.
[0145] A "generative AI model" is an artificial intelligence model that learns from collected data and has the ability to make inferences and predictions about new data.
[0146] A "unified space" refers to a data space that centrally aggregates data obtained from different sources and perspectives and presents it in a unified format and perspective.
[0147] A "clustering algorithm" is an algorithm used to group data based on similarity and calculate the relative positions of viewpoints.
[0148] "Request" means a request or inquiry sent by a User to a Server for specific information.
[0149] "Risk assessment" is the process of using collected data and trained generative AI models to assess the risk of potential cybersecurity-related threats and attacks.
[0150] "Real-time alerts" refer to warning messages that immediately notify users when cybersecurity threats increase and encourage them to take prompt action.
[0151] This invention is a system that trains a generative AI model using data collected from public sources and provides information from a unified perspective. The system functions through the following stages: [collection of public information], [data preprocessing], [storage in a database], [training of a generative AI model], [mapping to a unified space], [issuance of risk assessment and real-time alerts], and [provision of a user interface].
[0152] Hardware and Software Configuration
[0153] Hardware:
[0154] Server: Responsible for data collection, storage, model training, inference, and request processing.
[0155] Devices: Smartphones, smart glasses, head-mounted displays, robots, etc. Serve as the interface with the user.
[0156] software:
[0157] Data collection: API, web scraping tools (BeautifulSoup, Selenium)
[0158] Data preprocessing: Python (Pandas, NumPy)
[0159] Model training: TensorFlow, PyTorch
[0160] Clustering: Scikit-learn
[0161] Natural Language Processing: SpaCy, NLTK
[0162] User interface: Flutter (smartphones), Unity (smart glasses, head-mounted displays, robots)
[0163] Data collection and preprocessing
[0164] The server uses APIs and web scraping tools to collect data from public sources on the Internet (government databases, news sites, social media, etc.). The collected data is pre-processed using Python libraries to clean and standardize the format. The pre-processed data is then structured and stored in a database.
[0165] Generative AI models and mapping to a unified space
[0166] The saved data is input into a generative AI model, which is then trained. The generative AI model is developed using TensorFlow and PyTorch. The trained generative AI model maps the data obtained from multiple viewpoints into a unified space. The relative positions of the viewpoints are calculated and visualized using Scikit-learn's clustering algorithm.
[0167] Risk Assessment and Real-Time Alerts
[0168] The server evaluates cybersecurity-related risks based on the generated information and sends real-time alerts to users' systems about high-risk threats, allowing them to take prompt action.
[0169] User Interface
[0170] User requests are sent to the server via the device. The request is analyzed and relevant information is searched for and generated. User interfaces developed using Flutter and Unity are provided to users via smartphones, smart glasses, head-mounted displays, and robots.
[0171] Examples of concrete examples and prompts
[0172] For example, if a new ransomware attack is reported, the server collects relevant information, which is then scrutinized by the generative AI model and sent to the device, where it is notified to the user visually and audibly, providing immediate countermeasures.
[0173] Prompt Sentence Examples
[0174] User: "What's the latest ransomware attack?"
[0175] Security Assistant AI: "The latest ransomware attack is called ____ and occurred on ____ / ____ / __. Details of the attack are as follows: ____. As a countermeasure, we recommend ____."
[0176] This allows users to respond quickly to the latest cyber threats and strengthen their security.
[0177] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0178] Step 1:
[0179] The server uses APIs and web scraping tools (BeautifulSoup, Selenium, etc.) to collect data from public sources (government databases, news sites, social media, etc.). The input of this step is the public source, and the output is the collected raw data. Specifically, the server accesses the specified URL and retrieves the required information.
[0180] Step 2:
[0181] The server performs preprocessing on the collected raw data. Preprocessing includes data cleaning (removing unnecessary data and special characters), formatting unification (e.g., changing the date and time format), and text standardization (e.g., lowercasing). The input of this step is the collected raw data, and the output is preprocessed clean data. Specifically, the data is converted using Python's Pandas and NumPy libraries.
[0182] Step 3:
[0183] The server saves the preprocessed data in a database. The input of this step is clean data, and the output is data stored in the database. Specifically, it inserts and updates data in the database using SQL queries.
[0184] Step 4:
[0185] The server inputs the clean data into the generative AI model and trains the model. The input for this step is clean data retrieved from the database, and the output is a trained generative AI model. Specifically, the generative AI model is trained using TensorFlow or PyTorch.
[0186] Step 5:
[0187] The server uses a trained generative AI model to map multiple viewpoints into a unified space and calculates the relative positions of the viewpoints using Scikit-learn's clustering algorithm. The input to this step is the generative AI model and clean data, and the output is viewpoint data mapped into a unified space. Specifically, the clustering algorithm is applied to group the data.
[0188] Step 6:
[0189] The user sends an information request to the server through their device. The input in this step is the request from the user, and the output is the analysis result of the request content. Specifically, the user inputs the request using a smartphone or smart glasses.
[0190] Step 7:
[0191] Based on the received request, the server analyzes the request content using natural language processing (SpaCy or NLTK) and searches for related information from a unified space. The input to this step is the analysis result, and the output is the searched related information. Specifically, it extracts related information from a database based on the request content.
[0192] Step 8:
[0193] The server generates information using a generative AI model based on the searched related information and sends it to the terminal. The input of this step is the searched related information, and the output is the generated information. Specifically, the server generates information based on the generative AI model and sends it to the terminal.
[0194] Step 9:
[0195] The terminal provides the generated information received from the server to the user. The input of this step is the generated information from the server, and the output is information displayed and notified in a format that the user can understand. Specific operations include providing information to the user visually or audibly using a smartphone, smart glasses, a head-mounted display, or a robot.
[0196] Through this series of processing steps, the system alerts users in real time to high-risk cyber threats, enabling them to respond quickly.
[0197] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0198] This invention is a system that trains a generative AI model using data collected from public sources and provides information from a unified perspective. This system also incorporates an emotion engine that recognizes user emotions, improving the effectiveness of information provision. The system consists of a server, a terminal, and a user.
[0199] Public information collection and preprocessing
[0200] First, the server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, it obtains the necessary data using APIs and web scraping technology. For example, it collects the latest policy information on COVID-19 from the government's open data portal and related articles from news sites. Next, the server preprocesses the collected data. It cleans the data, standardizes the format, removes special characters and spaces, and arranges the data into a unified format.
[0201] Storing in a database and training the generated AI model
[0202] The preprocessed data is stored in a database by the server. The server then retrieves the latest data from the database and inputs it into the generative AI model. This generative AI model has the ability to learn from the collected data and make inferences and predictions about new data. The server trains the generative AI model so that it always reflects the latest information.
[0203] Mapping to a unified space and calculating the relative viewpoints
[0204] The server runs a technology that maps multiple viewpoints into a unified space using a trained generative AI model. During this process, data from different sources and viewpoints is centrally aggregated. Furthermore, a clustering algorithm is used to calculate and visualize the relative positions of each viewpoint. As a result, users can easily grasp the relevance and reliability of the information.
[0205] Emotion engine recognizes user emotions
[0206] The device sends user input and behavioral data (e.g., search history, clicks, input text, etc.) to the emotion engine, which analyzes the data and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.). This analysis is performed using natural language processing (NLP) and machine learning algorithms.
[0207] Receiving user requests and generating information
[0208] A user requests specific information through a device (such as a PC or smartphone). For example, they send a request such as, "Please tell me the latest measures against COVID-19." The device then sends this request to the server. The server receives the request from the user and analyzes it using natural language processing (NLP). Based on the analyzed request, the server searches for relevant information from a unified space and generates appropriate comments or policy proposals using a generative AI model.
[0209] Customize information based on user sentiment
[0210] Based on the analysis results from the emotion engine, the server adjusts the tone and content of the information generated by the generative AI model. For example, if the user is feeling anxious, the information provided will be generated in a reassuring tone. In this way, information is customized according to the user's emotional state, making it possible to provide more effective information.
[0211] Providing information and concrete examples
[0212] The server sends the generated information to the device, which then displays the results to the user. For example, if a user requests, "I want to know about COVID-19 countermeasures," the server generates the latest policy information and suggestions, and based on the results of the emotion engine, sends information such as, "The latest government policy recommends vaccination in the XX region. In addition, the self-isolation period has been changed from XX to XX days. If you are concerned, please refer to the detailed guidelines." to the device, which then displays this to the user. This system allows users to easily access the latest and most reliable information, which can be used to understand policies and support their daily lives. Furthermore, providing information based on the user's emotions increases user satisfaction and trust.
[0213] The processing flow will be explained below.
[0214] Step 1:
[0215] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.) by using REST APIs to retrieve data in JSON format or by using web scraping technology to obtain the required information.
[0216] Step 2:
[0217] The server pre-processes the data it receives, cleaning it (for example, removing whitespace and special characters), standardizing date and number formats, and arranging the data into a consistent format.
[0218] Step 3:
[0219] The server saves the pre-processed data to a database, where the information is stored efficiently and securely using INSERT operations into an SQL database.
[0220] Step 4:
[0221] The server retrieves the latest data from the database and inputs it into the generative AI model, which then uses the newly collected data to train the generative AI model.
[0222] Step 5:
[0223] The server trains the generative AI model, running machine learning algorithms that enable the model to make accurate predictions and inferences on new data.
[0224] Step 6:
[0225] The server maps multiple perspectives from the generative AI model into a unified space. Specifically, it plots data from different sources and perspectives in a multidimensional space and calculates their spatial relationships using a clustering algorithm.
[0226] Step 7:
[0227] The terminal sends the user's input data and behavioral data (e.g., past search history, click status, entered text, etc.) to the emotion engine.
[0228] Step 8:
[0229] The emotion engine analyzes data sent from the device and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.) using natural language processing (NLP) and machine learning algorithms.
[0230] Step 9:
[0231] Users request specific information through their device, for example, by typing in a request such as "Please tell me the latest information on COVID-19 countermeasures" and submitting it.
[0232] Step 10:
[0233] The terminal sends a request from the user to the server, and HTTP requests are generally used as the communication protocol.
[0234] Step 11:
[0235] The server analyzes the request received from the terminal, and uses natural language processing (NLP) to analyze the content of the request and identify its requirements.
[0236] Step 12:
[0237] The server searches for relevant information from the unified space based on the analysis results, leveraging the capabilities of SQL queries and generative AI models to extract the necessary information.
[0238] Step 13:
[0239] The server uses generative AI models to generate comments and policy proposals based on the retrieved information, using natural language generation (NLG) to provide the generated text in a format that is easy for users to understand.
[0240] Step 14:
[0241] The server adjusts the tone and content of the generated information based on the user's emotion analysis results obtained from the emotion engine. For example, if the user is feeling anxious, the information provided will be generated in a reassuring tone.
[0242] Step 15:
[0243] The server then sends the generated information to the terminal. Again, HTTP responses are generally used as the communication protocol.
[0244] Step 16:
[0245] The device displays the information received from the server to the user, either by updating a web page or displaying the information in the mobile app UI, making it easily accessible to the user.
[0246] This series of steps allows users to efficiently utilize public information and obtain the information they need quickly and accurately, as well as receive information tailored to their own emotions.
[0247] Example 2
[0248] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0249] In conventional information provision systems, it has been difficult to collect data from multiple information sources and provide users with reliable information from a unified perspective. Furthermore, the information provided does not take into account the user's emotional state, resulting in a decrease in user satisfaction. The present invention aims to solve these problems by handling data collected from a variety of information sources in a unified manner and providing information that corresponds to the user's emotional state.
[0250] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0251] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a database, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple perspectives into a unified space from the trained generative AI model, means for analyzing emotions using user input data and behavioral data, means for receiving requests from users, means for searching and generating information based on the requests, means for providing the generated information to the user, and means for adjusting the tone and content of the generated information based on the results of the emotion analysis. This makes it possible to handle data collected from multiple information sources in a unified manner and provide appropriate information according to the user's emotional state.
[0252] "Public information sources" refer to information sources such as databases, news sites, and social media that are publicly available on the Internet.
[0253] "Means of collecting data" refers to the processes and techniques used to obtain the required information from public sources using APIs and web scraping techniques.
[0254] "Preprocessing" refers to the process of removing noise from collected data, standardizing the format, removing special characters and unnecessary spaces, and otherwise preparing the data for easier analysis.
[0255] "Means of database storage" refers to the process or technique by which pre-processed data is stored in a data management system.
[0256] A "generative AI model" refers to an artificial intelligence model that has the ability to learn from large amounts of data and make inferences and predictions based on new data.
[0257] "Training means" refers to the process or technique of inputting data into a generative AI model and optimizing the model's parameters.
[0258] "Means of mapping to a unified space" refers to the process of centrally aggregating data from multiple perspectives obtained by a generative AI model and visualizing the relationships between the perspectives in a low-dimensional space.
[0259] "Means of emotion analysis" refers to natural language processing and machine learning technologies that recognize users' emotional states based on their input data and behavioral data.
[0260] "Means for receiving requests from users" refers to the process by which the server receives inquiries or information requests sent by users through their terminals.
[0261] "Means for information retrieval and generation" refers to the process of analyzing a user's request, searching for relevant information from a database based on that request, and providing appropriate information using a generative AI model.
[0262] "Means for providing information to a user" refers to the process of displaying the generated information to a user through a terminal.
[0263] "Measures to adjust tone and content" refers to the process of changing the expression and content of generated information based on the results of sentiment analysis in accordance with the user's emotional state.
[0264] This system trains a generative AI model using data collected from public sources to provide information from a unified perspective. It also incorporates an emotion engine that recognizes the user's emotional state and provides information accordingly. The system consists of a server, a terminal, and a user.
[0265] Public information collection and preprocessing
[0266] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, the server obtains data using APIs and web scraping technologies (e.g., BeautifulSoup or Scrapy). For example, the server collects the latest policy information about COVID-19 from the government's open data portal and related articles from news sites. The server then preprocesses the collected data. This involves removing noise, standardizing the format, and removing special characters and spaces, thereby arranging the data into a unified format. This makes it easier to input into the generative AI model.
[0267] Storing in a database and training the generated AI model
[0268] The preprocessed data is stored in a database by the server. Specifically, the server connects to a database system such as MySQL or MongoDB and stores the data. The server then retrieves the latest data from the database and inputs it into a generative AI model (e.g., GPT-3 or BERT). This generative AI model has the ability to learn from the collected data and make inferences and predictions based on new data. The server trains the generative AI model using libraries such as TensorFlow and PyTorch to ensure that it always reflects the latest information.
[0269] Mapping information to a unified space and calculating spatial relationships
[0270] The server maps multiple perspectives obtained from the trained generative AI model into a unified space. In this process, data obtained from different information sources and perspectives is centrally aggregated. The server then uses a clustering algorithm (e.g., k-means) to calculate the relative positions of each perspective and visualize them in a low-dimensional space. This allows users to easily grasp the relevance and reliability of the information.
[0271] Emotion engine recognizes user emotions
[0272] The device sends the user's input data and behavioral data (e.g., search history, clicks, input text, etc.) to the emotion engine. The emotion engine analyzes this data and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.). Specifically, it performs the analysis using natural language processing (NLP) and machine learning algorithms (e.g., the emotion analysis model VADER).
[0273] Receiving user requests and generating information
[0274] A user requests specific information through a device (e.g., a PC or smartphone). For example, the user types, "Please tell me the latest measures against COVID-19." This request is sent from the device to the server. The server receives the request and analyzes it using natural language processing (NLP) technology. Based on the analysis results, the server searches for relevant information from a unified space and generates appropriate information using a generative AI model.
[0275] Regulating and providing information based on emotions
[0276] Based on the analysis results from the emotion engine, the server adjusts the tone and content of the generated information. For example, if the user is feeling anxious, the server will provide information in a reassuring tone, such as, "Don't worry, the latest government policy recommends vaccination in the XX area." The final information is sent from the server to the device, which then displays it to the user. This allows users to always have easy access to the latest and most reliable information, allowing them to use the information with greater peace of mind.
[0277] Examples of concrete examples and prompts
[0278] For example, if you receive a request such as "I want to know the latest information on COVID-19 countermeasures," the prompt might look like this:
[0279] What is the latest information on COVID-19 prevention measures? My current emotional state is anxiety.
[0280] By inputting this prompt into the generative AI model, appropriate information is provided.
[0281] This system provides more effective information provision by customizing information according to the user's emotional state, and by centralizing data collected from various sources, users can quickly obtain reliable information.
[0282] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0283] Step 1:
[0284] The server collects data from public sources (government databases, news sites, social media, etc.). Specifically, the server obtains information using APIs or web scraping technologies (e.g., BeautifulSoup or Scrapy). The input is the URL or API key of the public source, and the output is the collected raw data. For example, the server obtains the latest policy information about COVID-19 from the government's open data portal.
[0285] Step 2:
[0286] The server preprocesses the collected data. The input is the raw data collected in step 1, and the output is clean, uniformly formatted data. Specifically, the server uses regular expressions to remove noise, standardize the format, and remove special characters and whitespace. This prepares the data in a form that is easy to input into the generative AI model.
[0287] Step 3:
[0288] The server saves the preprocessed data in a database. The input is the preprocessed data, and the output is the data stored in the database. Specifically, the server connects to a database system such as MySQL or MongoDB and executes SQL statements to insert data. This allows the data to be managed in an organized manner.
[0289] Step 4:
[0290] The server retrieves the latest data from the database and inputs it into the generative AI model. The input is the formatted data retrieved from the database, and the output is the data processed by the generative AI model. Specifically, the server retrieves the data and inputs it into the generative AI model (e.g., GPT-3 or BERT) using libraries such as TensorFlow or PyTorch.
[0291] Step 5:
[0292] The server trains the generative AI model. The input is the training data and the model's initial parameters, and the output is the trained model. For example, the server uses a dataset to optimize the model's parameters and improve its inference ability on new data. This training process is performed using libraries such as TensorFlow and PyTorch.
[0293] Step 6:
[0294] The server maps multiple viewpoints from a trained generative AI model into a unified space. The input is viewpoint data generated by the model, and the output is a unified viewpoint space. Specifically, the server visualizes the viewpoint data in a low-dimensional space using UMAP (Uniform Manifold Approximation and Projection) or a clustering algorithm (e.g., k-means).
[0295] Step 7:
[0296] The device sends the user's input data and behavioral data to the emotion engine. The input is the user's input data (search history, clicks, input text, etc.), and the output is the emotion analysis results. The device sends this data to the emotion engine, which then analyzes it using natural language processing technology.
[0297] Step 8:
[0298] A user requests specific information through a device. The input is the user's request (e.g., "Please tell me the latest measures against COVID-19"), and the output is the request data. This request is sent from the device to the server.
[0299] Step 9:
[0300] The server receives the user's request and analyzes it using natural language processing technology. The input is the text data of the user's request, and the output is the analysis result. Based on the analysis result, the server searches for relevant information from the unified space and generates appropriate information using a generative AI model.
[0301] Step 10:
[0302] The server adjusts the tone and content of the generated information based on the analysis results from the emotion engine. The input is the emotion analysis result and the generated information, and the output is the adjusted information. For example, if the user is feeling anxious, the server will provide information in a reassuring tone.
[0303] Step 11:
[0304] The server sends the generated information to the terminal, which then displays it to the user. The input is the adjusted information, and the output is the information displayed to the user. The terminal displays on the user's smartphone or PC, "The latest government policy recommends vaccination in the XX area. Additionally, the home quarantine period has been changed from XX to XX days."
[0305] (Application example 2)
[0306] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0307] In modern society, the importance of security information is increasing day by day. However, there is still a lack of systems that allow users to obtain reliable security information in real time and provide that information tailored to the user's emotional state. Furthermore, there is the challenge of properly recognizing users' emotions, such as anxiety and fear, and providing customized information based on those emotions. Therefore, it is necessary to provide a system that allows users to receive security information with peace of mind.
[0308] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0309] In this invention, the server includes means for collecting data from public information sources, means for pre-processing the collected data, and means for storing the pre-processed data in a database, thereby enabling security information collected from public information sources to be provided to the user in real time and further customized by recognizing the user's emotional state.
[0310] "Public sources" are sources that are free or publicly available on the internet and can be accessed by anyone. Examples include government databases, news sites, and social media.
[0311] "Methods of data collection" refers to technologies and tools used to automatically obtain the required data from public sources on the Internet, including the use of APIs and web scraping techniques.
[0312] "Preprocessing methods" refers to techniques for improving the quality of collected data and arranging it into a unified format, such as data cleaning, formatting standardization, and removal of special characters.
[0313] "Means of storing data in a database" refers to a system that efficiently stores pre-processed data and manages it for easy retrieval and use. Examples include RDBMS and NoSQL databases.
[0314] "Means of input to generative AI models" refers to technologies that provide data stored in a database to generative AI models for use in training and prediction.
[0315] "Training methods" refers to the techniques used to train generative AI models with the latest data and improve their performance, including machine learning algorithms and data feeds.
[0316] "Means of mapping to a unified space" refers to techniques that unify data from different sources and perspectives to allow users to understand the relevance and reliability of the information. Specifically, this includes dimensionality reduction and clustering algorithms.
[0317] "Means for receiving requests" refers to the technology used to obtain information requests from users and communicate them to the system, including the user interface and the NLP engine.
[0318] "Information search and generation means" refers to technologies that search for relevant information from databases based on user requests and generate appropriate answers or comments, including search algorithms and generative AI models.
[0319] "Means for providing to users" refers to technologies for presenting generated information to users in an easy-to-understand manner, including user interfaces and notification systems.
[0320] "Emotional state recognition" refers to technologies for identifying a user's current emotional state based on their input and behavioral data, including natural language processing and machine learning algorithms.
[0321] "Information tailoring measures" refers to techniques for adjusting the tone and content of information provided in response to the perceived emotional state of the user, including tone modification and information filtering.
[0322] This invention is a system that collects data from public sources, uses a generative AI model to generate information requested by users, and further customizes the information according to the user's emotional state. The system mainly consists of three elements: a server, a terminal, and a user.
[0323] System configuration
[0324] server
[0325] The server has the following main functions:
[0326] Data collection from public sources: The server collects data in real time from public sources on the Internet, using APIs and web scraping techniques to extract necessary security information from government databases, news sites, social media, etc.
[0327] Data preprocessing: The collected data is preprocessed by data cleaning, standardizing the format, removing special characters, etc.
[0328] Database storage: The preprocessed data is stored in a database that can be efficiently searched and updated, typically using an RDBMS or NoSQL database.
[0329] Information generation using a generative AI model: The stored data is input into a generative AI model to train the model. The trained generative AI model is then used to generate information based on user requests.
[0330] Emotion recognition: Analyzes user input data (e.g., search history, clicks, input text) and uses an emotion engine to recognize the user's current emotional state.
[0331] Information customization: Adjusting the tone and content of generated information based on perceived emotional state.
[0332] Terminal
[0333] The device (e.g., smartphone or PC) has the following features:
[0334] Sending a user request: A user requests specific information and sends it to the server via their device. For example, a request could be, "Please tell me about the latest cyber-attack information."
[0335] Receiving and displaying information: Receives information generated by the server and displays it to the user. The display method can be in various formats such as text or notification.
[0336] user
[0337] The user does the following:
[0338] Information Request: Using a smartphone or PC, a user requests specific information. This request is made in natural language and sent to the server.
[0339] Review the information: Review the information received and take further action as necessary.
[0340] Specific examples of processing
[0341] When a user requests, "I'd like to know about recent cyber-attack trends," the server first collects the latest security-related data from public sources. This data is preprocessed and stored in a database. Then, a generative AI model is used to generate appropriate information and customize it based on the user's emotional state. Finally, the generated, specific answer (e.g., "Recent cyber-attacks have been characterized by large-scale phishing campaigns and the emergence of new ransomware. In particular, there has been an increase in techniques that identify and encrypt sensitive data within companies and demand ransom. Effective defensive measures include regular software updates and strengthened employee education. If you have concerns, please refer to our more detailed security policy") is provided to the user via their device.
[0342] Prompt Sentence Examples
[0343] Examples of specific prompts based on user input include:
[0344] User input: "I'd like to know about recent trends in cyber attacks."
[0345] Prompt: "User's feelings: anxious. Collect the security-related information: (Recently collected security information). User's request: I would like to know about recent cyber-attack trends."
[0346] This specific example allows users to obtain the latest and most reliable security information in real time, while also providing specific measures to alleviate their concerns.
[0347] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0348] Step 1:
[0349] The server collects data from public sources. Specifically, it uses APIs and web scraping technology to obtain real-time data from government databases, news sites, and social media on the Internet. The collected data is temporarily stored as raw data. Inputs include the URL and API key of the target public source, and the obtained raw data is included as output.
[0350] Step 2:
[0351] The server preprocesses the collected data. Specific operations include data cleaning (removing missing values and outliers), standardizing formats (e.g., standardizing character codes and date formats), removing special characters, etc. The input data is the raw data collected in the previous step, and the output data is the preprocessed clean data.
[0352] Step 3:
[0353] The server stores the preprocessed data in a database. Specifically, it uses an RDBMS or NoSQL database to efficiently store data. The input data is the preprocessed clean data, and the output is the storage data stored in the database.
[0354] Step 4:
[0355] The server inputs the stored data into the generative AI model and trains the model. It retrieves the pre-processed and stored data from the database and supplies it to the generative AI model, updating the model with the latest information and training it. The input data is the data retrieved from the database, and the output data is the trained generative AI model.
[0356] Step 5:
[0357] The server maps multiple perspectives from the trained generative AI model into a unified space. Specifically, it aggregates data from different information sources and perspectives into a unified space using a clustering algorithm and visualizes it. The input data is the trained generative AI model, and the output data is the mapped perspective data.
[0358] Step 6:
[0359] The server receives requests from users and analyzes them. When a user sends a request from a device, the server analyzes the request using natural language processing (NLP). The input data is the user's request, and the output data is the analysis result.
[0360] Step 7:
[0361] The server searches and generates information based on the analyzed request, searches for relevant information from the database, and generates an appropriate answer using a generative AI model. The input data is the analysis result and the storage data saved in the database, and the output data is the generated answer.
[0362] Step 8:
[0363] The server recognizes the user's emotional state. It uses an emotion engine to analyze the user's emotional state (e.g., relief, anxiety, interest, etc.) from their search history and input text. The input data is the user's behavioral data, and the output data is the recognized emotional state.
[0364] Step 9:
[0365] The server adjusts the generated information based on the recognized emotional state, specifically by changing the tone of the text or filtering the information. The input data is the recognized emotional state and the generated response, and the output data is the adjusted response.
[0366] Step 10:
[0367] The terminal provides the generated information received from the server to the user. The terminal presents the information to the user in text format, notification format, etc. The input data is the adjusted response from the server, and the output data is the information received by the user.
[0368] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0369] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0370] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0371] [Second embodiment]
[0372] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0373] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0374] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0375] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0376] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0377] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0378] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0379] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0380] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0381] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0382] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0383] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0384] The present invention is a system that trains a generative AI model using data collected from public information sources and provides information from a unified perspective. This system consists of a server, a terminal, and a user.
[0385] Public information collection and preprocessing
[0386] First, the server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, it obtains the necessary data using APIs and web scraping technology. For example, it collects the latest policy information on COVID-19 from the government's open data portal and related articles from news sites. Next, the server preprocesses the collected data. It cleans the data, standardizes the format, removes special characters and spaces, and arranges the data into a unified format.
[0387] Storing in a database and training the generated AI model
[0388] The preprocessed data is stored in a database by the server. The server then retrieves the latest data from the database and inputs it into the generative AI model. This generative AI model has the ability to learn from the collected data and make inferences and predictions about new data. The server trains the generative AI model so that it always reflects the latest information.
[0389] Mapping to a unified space and calculating the relative viewpoints
[0390] The server runs a technology that maps multiple viewpoints into a unified space using a trained generative AI model. During this process, data from different sources and viewpoints is centrally aggregated. Furthermore, a clustering algorithm is used to calculate and visualize the relative positions of each viewpoint. As a result, users can easily grasp the relevance and reliability of the information.
[0391] Receiving user requests and generating information
[0392] A user requests specific information through a device (such as a PC or smartphone). For example, they send a request such as, "Please tell me the latest measures against COVID-19." The device then sends this request to the server. The server receives the request from the user and analyzes it using natural language processing (NLP). Based on the analyzed request, the server searches for relevant information from a unified space and generates appropriate comments or policy proposals using a generative AI model.
[0393] Providing information and concrete examples
[0394] The server sends the generated information to the device, which then displays the results to the user. For example, if a user requests, "I want to know about COVID-19 countermeasures," the server generates the latest policy information and suggestions, and sends information such as, "The latest government policy recommends vaccination in XX region. In addition, the home quarantine period has been changed from XX to XX days," to the device, which then displays this to the user. This system allows users to always easily access the latest, most reliable information, which can be used to understand policies and support their daily lives.
[0395] The processing flow will be explained below.
[0396] Step 1:
[0397] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.) by using REST APIs to retrieve data in JSON format or by using web scraping technology to obtain the required information.
[0398] Step 2:
[0399] The server pre-processes the data it receives, cleaning it (for example, removing whitespace and special characters), standardizing date and number formats, and arranging the data into a consistent format.
[0400] Step 3:
[0401] The server saves the pre-processed data to a database, where the information is stored efficiently and securely using INSERT operations into an SQL database.
[0402] Step 4:
[0403] The server retrieves the latest data from the database and inputs it into the generative AI model, which then uses the newly collected data to train the generative AI model.
[0404] Step 5:
[0405] The server trains the generative AI model, running machine learning algorithms that enable the model to make accurate predictions and inferences on new data.
[0406] Step 6:
[0407] The server maps multiple perspectives from the generative AI model into a unified space. Specifically, it plots data from different sources and perspectives in a multidimensional space and calculates their spatial relationships using a clustering algorithm.
[0408] Step 7:
[0409] The user requests specific information through the device, for example, "Please tell me the latest information on COVID-19 countermeasures," and sends it.
[0410] Step 8:
[0411] The terminal sends a request from the user to the server, and HTTP requests are generally used as the communication protocol.
[0412] Step 9:
[0413] The server analyzes the request received from the terminal, and uses natural language processing (NLP) to analyze the content of the request and identify its requirements.
[0414] Step 10:
[0415] The server searches for relevant information from the unified space based on the analysis results, leveraging the capabilities of SQL queries and generative AI models to extract the necessary information.
[0416] Step 11:
[0417] The server uses generative AI models to generate comments and policy proposals based on the retrieved information, using natural language generation (NLG) to provide the generated text in a format that is easy for users to understand.
[0418] Step 12:
[0419] The server then sends the generated information to the terminal. Again, HTTP responses are generally used as the communication protocol.
[0420] Step 13:
[0421] The device displays the information received from the server to the user, either by updating a web page or displaying the information in the mobile app UI, making it easily accessible to the user.
[0422] This series of steps allows users to efficiently utilize open public information and obtain the information they need quickly and accurately.
[0423] Example 1
[0424] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0425] In conventional information provision systems, when collecting data from public information sources and providing information from a unified perspective, multiple processes such as data preprocessing, training of generative AI models, and viewpoint mapping are required, and each process is often performed independently, resulting in problems that reduce the efficiency and accuracy of the entire system.In addition, there is a lack of functionality to properly calculate and visualize the positional relationships of viewpoints obtained from different information sources, making it difficult for users to quickly obtain reliable information.
[0426] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0427] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a data management device, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple viewpoints into a unified space from the trained generative AI model, means for calculating the positional relationship between viewpoints obtained from different information sources, means for receiving a request from a user, means for searching and generating information based on the request, and means for providing the generated information to the user, thereby enabling the user to quickly obtain reliable information and grasp a variety of information in a unified manner.
[0428] A "public information source" is a database or website that provides information that is publicly accessible on the Internet.
[0429] "Data preprocessing" is the process of cleaning, formatting, and removing unnecessary information from collected data to prepare it in a format that is easy to analyze and store.
[0430] A "data management device" refers to a database or storage system for efficiently and safely storing collected data.
[0431] A "generative artificial intelligence model" is a machine learning model that can learn from collected data and predict or generate data in the future.
[0432] "Training" is the learning process that a generative artificial intelligence model goes through to achieve optimal performance based on data.
[0433] A "unified space" is a data representation that centrally manages multiple perspectives collected from different sources and maps them in a form that makes comparison and analysis easy.
[0434] "Positional relationships between viewpoints" refers to information that indicates how multiple viewpoints are related to each other.
[0435] A "clustering algorithm" is a computational method for grouping data based on similarity.
[0436] "User Request" means a request for information entered by a user into the system.
[0437] "Information retrieval and generation means" refers to the process of locating appropriate information from databases and generative artificial intelligence models based on user requests and generating text, reports, etc. as needed.
[0438] "Generated Information" refers to information created by the system using a generative artificial intelligence model based on a user request.
[0439] The present invention is a system that trains a generative AI model using data collected from public information sources and provides information from a unified perspective. This system consists of a server, a terminal, and a user.
[0440] First, the server collects data from public sources on the internet, such as government databases, news sites, and social networking services. This process utilizes APIs and web scraping techniques, using tools such as Python's requests library and BeautifulSoup. For example, it obtains policy information about COVID-19 from government open data portals and collects related articles from news sites.
[0441] The server then preprocesses the collected data, which includes removing special characters and whitespace, cleaning, and formatting, using the Python pandas library to organize the data into a unified format.
[0442] The preprocessed data is stored in a data management device by the server. The server uses a database system such as MySQL or PostgreSQL to efficiently store the data. The stored data is then input into a generative artificial intelligence model to train the model. Machine learning frameworks such as TensorFlow and PyTorch are used for this training.
[0443] A trained generative artificial intelligence model maps multiple viewpoints into a unified space. The server integrates data from different sources and viewpoints and calculates the spatial relationships between viewpoints using a clustering algorithm in Python's Scikit-learn library. As a result, the relationships between viewpoints are visualized, and important information is centralized.
[0444] A user requests specific information through a device, such as a PC or smartphone. An example of a prompt would be "Please tell me the latest measures against COVID-19." The device then sends this request to the server, which analyzes it using natural language processing (NLP) tools. The server uses Python's spaCy or NLTK to analyze the request, searches for relevant information in a unified space based on the analysis results, and generates comments or policy proposals using a generative artificial intelligence model if necessary.
[0445] Finally, the server sends the generated information to the device, which then displays the results to the user. For example, information such as "The latest government policy recommends vaccination in the ____ region. Furthermore, the self-isolation period has been changed from ____ to ____ days" can be provided. In this way, users can quickly obtain the latest and most reliable information.
[0446] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0447] Step 1: Gather public information
[0448] The server collects data from public sources on the Internet, including government databases, news sites, social networking services, etc. The server uses Python's requests library to retrieve the required data from APIs, and also uses BeautifulSoup to scrape articles from news sites.
[0449] Input: URL or API endpoint of a public source
[0450] Output: Raw data (text, JSON, etc.)
[0451] Step 2: Preprocessing the data
[0452] The server preprocesses the collected data, which includes removing unnecessary whitespace and special characters, standardizing the format, etc. The server uses the Python pandas library to clean the data and convert it into a unified format.
[0453] Input: Raw data
[0454] Output: Preprocessed data (unified format)
[0455] Step 3: Saving to the database
[0456] The server stores the preprocessed data in a data management device (database), such as MySQL or PostgreSQL. The server establishes a connection to the database and inserts data using SQL queries.
[0457] Input: Preprocessed data
[0458] Output: Data stored in the database
[0459] Step 4: Training the generative AI model
[0460] The server retrieves the latest data from the database and trains a generative AI model using machine learning frameworks such as TensorFlow and PyTorch. The server inputs the data into the model and trains it over multiple epochs.
[0461] Input: Data retrieved from the database
[0462] Output: A trained generative AI model
[0463] Step 5: Mapping to a unified space
[0464] The server uses a trained generative AI model to map multiple viewpoints into a unified space using a clustering algorithm using Python's Scikit-learn library, which calculates the spatial relationships between viewpoints obtained from different sources.
[0465] Input: A trained generative AI model and associated data
[0466] Output: viewpoint data mapped to a unified space
[0467] Step 6: Calculate the viewpoint position
[0468] The server calculates the relative positions of the viewpoints mapped to the unified space and prepares for visualization. It calculates the distance and relationship between viewpoints and generates a network graph. This process uses the Python networkx library.
[0469] Input: Mapped viewpoint data
[0470] Output: Data for visualizing the relative positions of viewpoints
[0471] Step 7: Receiving a user request
[0472] The user requests specific information through the device, for example by entering a prompt such as "Please tell me the latest information on COVID-19 countermeasures." The device then sends this request to the server.
[0473] Input: User request (prompt sentence)
[0474] Output: Request sent to the server
[0475] Step 8: Generate information
[0476] The server analyzes requests received from users using natural language processing (NLP) tools, such as Python's spaCy and NLTK. Based on the analysis results, the server searches for relevant information in a unified space and generates appropriate comments and policy proposals using a generative AI model.
[0477] Input: The request sent to the server
[0478] Output: Generated information (comments and policy suggestions)
[0479] Step 9: Provide information
[0480] The server sends the generated information to the device, which then displays the results to the user, such as, "The latest government policy recommends vaccination in the ____ region. In addition, the home quarantine period has been changed from ____ to ____ days."
[0481] Input: Generated information
[0482] Output: Information displayed on the user's terminal
[0483] In this way, specific actions are performed at each step, allowing the user to quickly obtain reliable, up-to-date information.
[0484] (Application example 1)
[0485] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0486] In today's world, businesses and individuals are constantly exposed to the threat of cyberattacks, resulting in increased risk of serious economic losses and information leaks. Cyberthreats evolve daily, and information about them is provided fragmentedly from numerous sources, making it difficult to unify, understand, and respond immediately. Furthermore, the lack of a system that can assess risks in real time and issue immediate alerts for high-risk threats means that prompt responses are delayed, potentially resulting in greater damage. To address these challenges and improve the effectiveness of cybersecurity, a system is needed that can unify the latest information, quickly and accurately assess risks, and notify users in real time.
[0487] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0488] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a database, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple perspectives into a unified space from the trained generative AI model, means for receiving requests from users, means for searching and generating information based on the requests, means for providing the generated information to users, means for risk assessing cybersecurity-related information, and means for issuing alerts to users about high-risk threats in real time. This enables rapid and accurate risk assessment of cyberattacks and the issuance of alerts in real time, thereby strengthening the security measures of companies and individuals.
[0489] "Public sources" refer to data sources that are accessible to anyone and from which information can be obtained, such as government databases, news sites, and social media.
[0490] "Data preprocessing" is the process of cleaning collected data, standardizing the format, removing unnecessary special characters, etc., to improve the quality of the data.
[0491] "Database" refers to a collection of data that stores collected and preprocessed data and structures it for easy access and retrieval.
[0492] A "generative AI model" is an artificial intelligence model that learns from collected data and has the ability to make inferences and predictions about new data.
[0493] A "unified space" refers to a data space that centrally aggregates data obtained from different sources and perspectives and presents it in a unified format and perspective.
[0494] A "clustering algorithm" is an algorithm used to group data based on similarity and calculate the relative positions of viewpoints.
[0495] "Request" means a request or inquiry sent by a User to a Server for specific information.
[0496] "Risk assessment" is the process of using collected data and trained generative AI models to assess the risk of potential cybersecurity-related threats and attacks.
[0497] "Real-time alerts" refer to warning messages that immediately notify users when cybersecurity threats increase and encourage them to take prompt action.
[0498] This invention is a system that trains a generative AI model using data collected from public sources and provides information from a unified perspective. The system functions through the following stages: [collection of public information], [data preprocessing], [storage in a database], [training of a generative AI model], [mapping to a unified space], [issuance of risk assessment and real-time alerts], and [provision of a user interface].
[0499] Hardware and Software Configuration
[0500] Hardware:
[0501] Server: Responsible for data collection, storage, model training, inference, and request processing.
[0502] Devices: Smartphones, smart glasses, head-mounted displays, robots, etc. Serve as the interface with the user.
[0503] software:
[0504] Data collection: API, web scraping tools (BeautifulSoup, Selenium)
[0505] Data preprocessing: Python (Pandas, NumPy)
[0506] Model training: TensorFlow, PyTorch
[0507] Clustering: Scikit-learn
[0508] Natural Language Processing: SpaCy, NLTK
[0509] User interface: Flutter (smartphones), Unity (smart glasses, head-mounted displays, robots)
[0510] Data collection and preprocessing
[0511] The server uses APIs and web scraping tools to collect data from public sources on the Internet (government databases, news sites, social media, etc.). The collected data is pre-processed using Python libraries to clean and standardize the format. The pre-processed data is then structured and stored in a database.
[0512] Generative AI models and mapping to a unified space
[0513] The saved data is input into a generative AI model, which is then trained. The generative AI model is developed using TensorFlow and PyTorch. The trained generative AI model maps the data obtained from multiple viewpoints into a unified space. The relative positions of the viewpoints are calculated and visualized using Scikit-learn's clustering algorithm.
[0514] Risk Assessment and Real-Time Alerts
[0515] The server evaluates cybersecurity-related risks based on the generated information and sends real-time alerts to users' systems about high-risk threats, allowing them to take prompt action.
[0516] User Interface
[0517] User requests are sent to the server via the device. The request is analyzed and relevant information is searched for and generated. User interfaces developed using Flutter and Unity are provided to users via smartphones, smart glasses, head-mounted displays, and robots.
[0518] Examples of concrete examples and prompts
[0519] For example, if a new ransomware attack is reported, the server collects relevant information, which is then scrutinized by the generative AI model and sent to the device, where it is notified to the user visually and audibly, providing immediate countermeasures.
[0520] Prompt Sentence Examples
[0521] User: "What's the latest ransomware attack?"
[0522] Security Assistant AI: "The latest ransomware attack is called ____ and occurred on ____ / ____ / __. Details of the attack are as follows: ____. As a countermeasure, we recommend ____."
[0523] This allows users to respond quickly to the latest cyber threats and strengthen their security.
[0524] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0525] Step 1:
[0526] The server uses APIs and web scraping tools (BeautifulSoup, Selenium, etc.) to collect data from public sources (government databases, news sites, social media, etc.). The input of this step is the public source, and the output is the collected raw data. Specifically, the server accesses the specified URL and retrieves the required information.
[0527] Step 2:
[0528] The server performs preprocessing on the collected raw data. Preprocessing includes data cleaning (removing unnecessary data and special characters), formatting unification (e.g., changing the date and time format), and text standardization (e.g., lowercasing). The input of this step is the collected raw data, and the output is preprocessed clean data. Specifically, the data is converted using Python's Pandas and NumPy libraries.
[0529] Step 3:
[0530] The server saves the preprocessed data in a database. The input of this step is clean data, and the output is data stored in the database. Specifically, it inserts and updates data in the database using SQL queries.
[0531] Step 4:
[0532] The server inputs the clean data into the generative AI model and trains the model. The input for this step is clean data retrieved from the database, and the output is a trained generative AI model. Specifically, the generative AI model is trained using TensorFlow or PyTorch.
[0533] Step 5:
[0534] The server uses a trained generative AI model to map multiple viewpoints into a unified space and calculates the relative positions of the viewpoints using Scikit-learn's clustering algorithm. The input to this step is the generative AI model and clean data, and the output is viewpoint data mapped into a unified space. Specifically, the clustering algorithm is applied to group the data.
[0535] Step 6:
[0536] The user sends an information request to the server through their device. The input in this step is the request from the user, and the output is the analysis result of the request content. Specifically, the user inputs the request using a smartphone or smart glasses.
[0537] Step 7:
[0538] Based on the received request, the server analyzes the request content using natural language processing (SpaCy or NLTK) and searches for related information from a unified space. The input to this step is the analysis result, and the output is the searched related information. Specifically, it extracts related information from a database based on the request content.
[0539] Step 8:
[0540] The server generates information using a generative AI model based on the searched related information and sends it to the terminal. The input of this step is the searched related information, and the output is the generated information. Specifically, the server generates information based on the generative AI model and sends it to the terminal.
[0541] Step 9:
[0542] The terminal provides the generated information received from the server to the user. The input of this step is the generated information from the server, and the output is information displayed and notified in a format that the user can understand. Specific operations include providing information to the user visually or audibly using a smartphone, smart glasses, a head-mounted display, or a robot.
[0543] Through this series of processing steps, the system alerts users in real time to high-risk cyber threats, enabling them to respond quickly.
[0544] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0545] This invention is a system that trains a generative AI model using data collected from public sources and provides information from a unified perspective. This system also incorporates an emotion engine that recognizes user emotions, improving the effectiveness of information provision. The system consists of a server, a terminal, and a user.
[0546] Public information collection and preprocessing
[0547] First, the server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, it obtains the necessary data using APIs and web scraping technology. For example, it collects the latest policy information on COVID-19 from the government's open data portal and related articles from news sites. Next, the server preprocesses the collected data. It cleans the data, standardizes the format, removes special characters and spaces, and arranges the data into a unified format.
[0548] Storing in a database and training the generated AI model
[0549] The preprocessed data is stored in a database by the server. The server then retrieves the latest data from the database and inputs it into the generative AI model. This generative AI model has the ability to learn from the collected data and make inferences and predictions about new data. The server trains the generative AI model so that it always reflects the latest information.
[0550] Mapping to a unified space and calculating the relative viewpoints
[0551] The server runs a technology that maps multiple viewpoints into a unified space using a trained generative AI model. During this process, data from different sources and viewpoints is centrally aggregated. Furthermore, a clustering algorithm is used to calculate and visualize the relative positions of each viewpoint. As a result, users can easily grasp the relevance and reliability of the information.
[0552] Emotion engine recognizes user emotions
[0553] The device sends user input and behavioral data (e.g., search history, clicks, input text, etc.) to the emotion engine, which analyzes the data and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.). This analysis is performed using natural language processing (NLP) and machine learning algorithms.
[0554] Receiving user requests and generating information
[0555] A user requests specific information through a device (such as a PC or smartphone). For example, they send a request such as, "Please tell me the latest measures against COVID-19." The device then sends this request to the server. The server receives the request from the user and analyzes it using natural language processing (NLP). Based on the analyzed request, the server searches for relevant information from a unified space and generates appropriate comments or policy proposals using a generative AI model.
[0556] Customize information based on user sentiment
[0557] Based on the analysis results from the emotion engine, the server adjusts the tone and content of the information generated by the generative AI model. For example, if the user is feeling anxious, the information provided will be generated in a reassuring tone. In this way, information is customized according to the user's emotional state, making it possible to provide more effective information.
[0558] Providing information and concrete examples
[0559] The server sends the generated information to the device, which then displays the results to the user. For example, if a user requests, "I want to know about COVID-19 countermeasures," the server generates the latest policy information and suggestions, and based on the results of the emotion engine, sends information such as, "The latest government policy recommends vaccination in the XX region. In addition, the self-isolation period has been changed from XX to XX days. If you are concerned, please refer to the detailed guidelines." to the device, which then displays this to the user. This system allows users to easily access the latest and most reliable information, which can be used to understand policies and support their daily lives. Furthermore, providing information based on the user's emotions increases user satisfaction and trust.
[0560] The processing flow will be explained below.
[0561] Step 1:
[0562] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.) by using REST APIs to retrieve data in JSON format or by using web scraping technology to obtain the required information.
[0563] Step 2:
[0564] The server pre-processes the data it receives, cleaning it (for example, removing whitespace and special characters), standardizing date and number formats, and arranging the data into a consistent format.
[0565] Step 3:
[0566] The server saves the pre-processed data to a database, where the information is stored efficiently and securely using INSERT operations into an SQL database.
[0567] Step 4:
[0568] The server retrieves the latest data from the database and inputs it into the generative AI model, which then uses the newly collected data to train the generative AI model.
[0569] Step 5:
[0570] The server trains the generative AI model, running machine learning algorithms that enable the model to make accurate predictions and inferences on new data.
[0571] Step 6:
[0572] The server maps multiple perspectives from the generative AI model into a unified space. Specifically, it plots data from different sources and perspectives in a multidimensional space and calculates their spatial relationships using a clustering algorithm.
[0573] Step 7:
[0574] The terminal sends the user's input data and behavioral data (e.g., past search history, click status, entered text, etc.) to the emotion engine.
[0575] Step 8:
[0576] The emotion engine analyzes data sent from the device and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.) using natural language processing (NLP) and machine learning algorithms.
[0577] Step 9:
[0578] Users request specific information through their device, for example, by typing in a request such as "Please tell me the latest information on COVID-19 countermeasures" and submitting it.
[0579] Step 10:
[0580] The terminal sends a request from the user to the server, and HTTP requests are generally used as the communication protocol.
[0581] Step 11:
[0582] The server analyzes the request received from the terminal, and uses natural language processing (NLP) to analyze the content of the request and identify its requirements.
[0583] Step 12:
[0584] The server searches for relevant information from the unified space based on the analysis results, leveraging the capabilities of SQL queries and generative AI models to extract the necessary information.
[0585] Step 13:
[0586] The server uses generative AI models to generate comments and policy proposals based on the retrieved information, using natural language generation (NLG) to provide the generated text in a format that is easy for users to understand.
[0587] Step 14:
[0588] The server adjusts the tone and content of the generated information based on the user's emotion analysis results obtained from the emotion engine. For example, if the user is feeling anxious, the information provided will be generated in a reassuring tone.
[0589] Step 15:
[0590] The server then sends the generated information to the terminal. Again, HTTP responses are generally used as the communication protocol.
[0591] Step 16:
[0592] The device displays the information received from the server to the user, either by updating a web page or displaying the information in the mobile app UI, making it easily accessible to the user.
[0593] This series of steps allows users to efficiently utilize public information and obtain the information they need quickly and accurately, as well as receive information tailored to their own emotions.
[0594] Example 2
[0595] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0596] In conventional information provision systems, it has been difficult to collect data from multiple information sources and provide users with reliable information from a unified perspective. Furthermore, the information provided does not take into account the user's emotional state, resulting in a decrease in user satisfaction. The present invention aims to solve these problems by handling data collected from a variety of information sources in a unified manner and providing information that corresponds to the user's emotional state.
[0597] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0598] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a database, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple perspectives into a unified space from the trained generative AI model, means for analyzing emotions using user input data and behavioral data, means for receiving requests from users, means for searching and generating information based on the requests, means for providing the generated information to the user, and means for adjusting the tone and content of the generated information based on the results of the emotion analysis. This makes it possible to handle data collected from multiple information sources in a unified manner and provide appropriate information according to the user's emotional state.
[0599] "Public information sources" refer to information sources such as databases, news sites, and social media that are publicly available on the Internet.
[0600] "Means of collecting data" refers to the processes and techniques used to obtain the required information from public sources using APIs and web scraping techniques.
[0601] "Preprocessing" refers to the process of removing noise from collected data, standardizing the format, removing special characters and unnecessary spaces, and otherwise preparing the data for easier analysis.
[0602] "Means of database storage" refers to the process or technique by which pre-processed data is stored in a data management system.
[0603] A "generative AI model" refers to an artificial intelligence model that has the ability to learn from large amounts of data and make inferences and predictions based on new data.
[0604] "Training means" refers to the process or technique of inputting data into a generative AI model and optimizing the model's parameters.
[0605] "Means of mapping to a unified space" refers to the process of centrally aggregating data from multiple perspectives obtained by a generative AI model and visualizing the relationships between the perspectives in a low-dimensional space.
[0606] "Means of emotion analysis" refers to natural language processing and machine learning technologies that recognize users' emotional states based on their input data and behavioral data.
[0607] "Means for receiving requests from users" refers to the process by which the server receives inquiries or information requests sent by users through their terminals.
[0608] "Means for information retrieval and generation" refers to the process of analyzing a user's request, searching for relevant information from a database based on that request, and providing appropriate information using a generative AI model.
[0609] "Means for providing information to a user" refers to the process of displaying the generated information to a user through a terminal.
[0610] "Measures to adjust tone and content" refers to the process of changing the expression and content of generated information based on the results of sentiment analysis in accordance with the user's emotional state.
[0611] This system trains a generative AI model using data collected from public sources to provide information from a unified perspective. It also incorporates an emotion engine that recognizes the user's emotional state and provides information accordingly. The system consists of a server, a terminal, and a user.
[0612] Public information collection and preprocessing
[0613] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, the server obtains data using APIs and web scraping technologies (e.g., BeautifulSoup or Scrapy). For example, the server collects the latest policy information about COVID-19 from the government's open data portal and related articles from news sites. The server then preprocesses the collected data. This involves removing noise, standardizing the format, and removing special characters and spaces, thereby arranging the data into a unified format. This makes it easier to input into the generative AI model.
[0614] Storing in a database and training the generated AI model
[0615] The preprocessed data is stored in a database by the server. Specifically, the server connects to a database system such as MySQL or MongoDB and stores the data. The server then retrieves the latest data from the database and inputs it into a generative AI model (e.g., GPT-3 or BERT). This generative AI model has the ability to learn from the collected data and make inferences and predictions based on new data. The server trains the generative AI model using libraries such as TensorFlow and PyTorch to ensure that it always reflects the latest information.
[0616] Mapping information to a unified space and calculating spatial relationships
[0617] The server maps multiple perspectives obtained from the trained generative AI model into a unified space. In this process, data obtained from different information sources and perspectives is centrally aggregated. The server then uses a clustering algorithm (e.g., k-means) to calculate the relative positions of each perspective and visualize them in a low-dimensional space. This allows users to easily grasp the relevance and reliability of the information.
[0618] Emotion engine recognizes user emotions
[0619] The device sends the user's input data and behavioral data (e.g., search history, clicks, input text, etc.) to the emotion engine. The emotion engine analyzes this data and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.). Specifically, it performs the analysis using natural language processing (NLP) and machine learning algorithms (e.g., the emotion analysis model VADER).
[0620] Receiving user requests and generating information
[0621] A user requests specific information through a device (e.g., a PC or smartphone). For example, the user types, "Please tell me the latest measures against COVID-19." This request is sent from the device to the server. The server receives the request and analyzes it using natural language processing (NLP) technology. Based on the analysis results, the server searches for relevant information from a unified space and generates appropriate information using a generative AI model.
[0622] Regulating and providing information based on emotions
[0623] Based on the analysis results from the emotion engine, the server adjusts the tone and content of the generated information. For example, if the user is feeling anxious, the server will provide information in a reassuring tone, such as, "Don't worry, the latest government policy recommends vaccination in the XX area." The final information is sent from the server to the device, which then displays it to the user. This allows users to always have easy access to the latest and most reliable information, allowing them to use the information with greater peace of mind.
[0624] Examples of concrete examples and prompts
[0625] For example, if you receive a request such as "I want to know the latest information on COVID-19 countermeasures," the prompt might look like this:
[0626] What is the latest information on COVID-19 prevention measures? My current emotional state is anxiety.
[0627] By inputting this prompt into the generative AI model, appropriate information is provided.
[0628] This system provides more effective information provision by customizing information according to the user's emotional state, and by centralizing data collected from various sources, users can quickly obtain reliable information.
[0629] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0630] Step 1:
[0631] The server collects data from public sources (government databases, news sites, social media, etc.). Specifically, the server obtains information using APIs or web scraping technologies (e.g., BeautifulSoup or Scrapy). The input is the URL or API key of the public source, and the output is the collected raw data. For example, the server obtains the latest policy information about COVID-19 from the government's open data portal.
[0632] Step 2:
[0633] The server preprocesses the collected data. The input is the raw data collected in step 1, and the output is clean, uniformly formatted data. Specifically, the server uses regular expressions to remove noise, standardize the format, and remove special characters and whitespace. This prepares the data in a form that is easy to input into the generative AI model.
[0634] Step 3:
[0635] The server saves the preprocessed data in a database. The input is the preprocessed data, and the output is the data stored in the database. Specifically, the server connects to a database system such as MySQL or MongoDB and executes SQL statements to insert data. This allows the data to be managed in an organized manner.
[0636] Step 4:
[0637] The server retrieves the latest data from the database and inputs it into the generative AI model. The input is the formatted data retrieved from the database, and the output is the data processed by the generative AI model. Specifically, the server retrieves the data and inputs it into the generative AI model (e.g., GPT-3 or BERT) using libraries such as TensorFlow or PyTorch.
[0638] Step 5:
[0639] The server trains the generative AI model. The input is the training data and the model's initial parameters, and the output is the trained model. For example, the server uses a dataset to optimize the model's parameters and improve its inference ability on new data. This training process is performed using libraries such as TensorFlow and PyTorch.
[0640] Step 6:
[0641] The server maps multiple viewpoints from a trained generative AI model into a unified space. The input is viewpoint data generated by the model, and the output is a unified viewpoint space. Specifically, the server visualizes the viewpoint data in a low-dimensional space using UMAP (Uniform Manifold Approximation and Projection) or a clustering algorithm (e.g., k-means).
[0642] Step 7:
[0643] The device sends the user's input data and behavioral data to the emotion engine. The input is the user's input data (search history, clicks, input text, etc.), and the output is the emotion analysis results. The device sends this data to the emotion engine, which then analyzes it using natural language processing technology.
[0644] Step 8:
[0645] A user requests specific information through a device. The input is the user's request (e.g., "Please tell me the latest measures against COVID-19"), and the output is the request data. This request is sent from the device to the server.
[0646] Step 9:
[0647] The server receives the user's request and analyzes it using natural language processing technology. The input is the text data of the user's request, and the output is the analysis result. Based on the analysis result, the server searches for relevant information from the unified space and generates appropriate information using a generative AI model.
[0648] Step 10:
[0649] The server adjusts the tone and content of the generated information based on the analysis results from the emotion engine. The input is the emotion analysis result and the generated information, and the output is the adjusted information. For example, if the user is feeling anxious, the server will provide information in a reassuring tone.
[0650] Step 11:
[0651] The server sends the generated information to the terminal, which then displays it to the user. The input is the adjusted information, and the output is the information displayed to the user. The terminal displays on the user's smartphone or PC, "The latest government policy recommends vaccination in the XX area. Additionally, the home quarantine period has been changed from XX to XX days."
[0652] (Application example 2)
[0653] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0654] In modern society, the importance of security information is increasing day by day. However, there is still a lack of systems that allow users to obtain reliable security information in real time and provide that information tailored to the user's emotional state. Furthermore, there is the challenge of properly recognizing users' emotions, such as anxiety and fear, and providing customized information based on those emotions. Therefore, it is necessary to provide a system that allows users to receive security information with peace of mind.
[0655] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0656] In this invention, the server includes means for collecting data from public information sources, means for pre-processing the collected data, and means for storing the pre-processed data in a database, thereby enabling security information collected from public information sources to be provided to the user in real time and further customized by recognizing the user's emotional state.
[0657] "Public sources" are sources that are free or publicly available on the internet and can be accessed by anyone. Examples include government databases, news sites, and social media.
[0658] "Methods of data collection" refers to technologies and tools used to automatically obtain the required data from public sources on the Internet, including the use of APIs and web scraping techniques.
[0659] "Preprocessing methods" refers to techniques for improving the quality of collected data and arranging it into a unified format, such as data cleaning, formatting standardization, and removal of special characters.
[0660] "Means of storing data in a database" refers to a system that efficiently stores pre-processed data and manages it for easy retrieval and use. Examples include RDBMS and NoSQL databases.
[0661] "Means of input to generative AI models" refers to technologies that provide data stored in a database to generative AI models for use in training and prediction.
[0662] "Training methods" refers to the techniques used to train generative AI models with the latest data and improve their performance, including machine learning algorithms and data feeds.
[0663] "Means of mapping to a unified space" refers to techniques that unify data from different sources and perspectives to allow users to understand the relevance and reliability of the information. Specifically, this includes dimensionality reduction and clustering algorithms.
[0664] "Means for receiving requests" refers to the technology used to obtain information requests from users and communicate them to the system, including the user interface and the NLP engine.
[0665] "Information search and generation means" refers to technologies that search for relevant information from databases based on user requests and generate appropriate answers or comments, including search algorithms and generative AI models.
[0666] "Means for providing to users" refers to technologies for presenting generated information to users in an easy-to-understand manner, including user interfaces and notification systems.
[0667] "Emotional state recognition" refers to technologies for identifying a user's current emotional state based on their input and behavioral data, including natural language processing and machine learning algorithms.
[0668] "Information tailoring measures" refers to techniques for adjusting the tone and content of information provided in response to the perceived emotional state of the user, including tone modification and information filtering.
[0669] This invention is a system that collects data from public sources, uses a generative AI model to generate information requested by users, and further customizes the information according to the user's emotional state. The system mainly consists of three elements: a server, a terminal, and a user.
[0670] System configuration
[0671] server
[0672] The server has the following main functions:
[0673] Data collection from public sources: The server collects data in real time from public sources on the Internet, using APIs and web scraping techniques to extract necessary security information from government databases, news sites, social media, etc.
[0674] Data preprocessing: The collected data is preprocessed by data cleaning, standardizing the format, removing special characters, etc.
[0675] Database storage: The preprocessed data is stored in a database that can be efficiently searched and updated, typically using an RDBMS or NoSQL database.
[0676] Information generation using a generative AI model: The stored data is input into a generative AI model to train the model. The trained generative AI model is then used to generate information based on user requests.
[0677] Emotion recognition: Analyzes user input data (e.g., search history, clicks, input text) and uses an emotion engine to recognize the user's current emotional state.
[0678] Information customization: Adjusting the tone and content of generated information based on perceived emotional state.
[0679] Terminal
[0680] The device (e.g., smartphone or PC) has the following features:
[0681] Sending a user request: A user requests specific information and sends it to the server via their device. For example, a request could be, "Please tell me about the latest cyber-attack information."
[0682] Receiving and displaying information: Receives information generated by the server and displays it to the user. The display method can be in various formats such as text or notification.
[0683] user
[0684] The user does the following:
[0685] Information Request: Using a smartphone or PC, a user requests specific information. This request is made in natural language and sent to the server.
[0686] Review the information: Review the information received and take further action as necessary.
[0687] Specific examples of processing
[0688] When a user requests, "I'd like to know about recent cyber-attack trends," the server first collects the latest security-related data from public sources. This data is preprocessed and stored in a database. Then, a generative AI model is used to generate appropriate information and customize it based on the user's emotional state. Finally, the generated, specific answer (e.g., "Recent cyber-attacks have been characterized by large-scale phishing campaigns and the emergence of new ransomware. In particular, there has been an increase in techniques that identify and encrypt sensitive data within companies and demand ransom. Effective defensive measures include regular software updates and strengthened employee education. If you have concerns, please refer to our more detailed security policy") is provided to the user via their device.
[0689] Prompt Sentence Examples
[0690] Examples of specific prompts based on user input include:
[0691] User input: "I'd like to know about recent trends in cyber attacks."
[0692] Prompt: "User's feelings: anxious. Collect the security-related information: (Recently collected security information). User's request: I would like to know about recent cyber-attack trends."
[0693] This specific example allows users to obtain the latest and most reliable security information in real time, while also providing specific measures to alleviate their concerns.
[0694] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0695] Step 1:
[0696] The server collects data from public sources. Specifically, it uses APIs and web scraping technology to obtain real-time data from government databases, news sites, and social media on the Internet. The collected data is temporarily stored as raw data. Inputs include the URL and API key of the target public source, and the obtained raw data is included as output.
[0697] Step 2:
[0698] The server preprocesses the collected data. Specific operations include data cleaning (removing missing values and outliers), standardizing formats (e.g., standardizing character codes and date formats), removing special characters, etc. The input data is the raw data collected in the previous step, and the output data is the preprocessed clean data.
[0699] Step 3:
[0700] The server stores the preprocessed data in a database. Specifically, it uses an RDBMS or NoSQL database to efficiently store data. The input data is the preprocessed clean data, and the output is the storage data stored in the database.
[0701] Step 4:
[0702] The server inputs the stored data into the generative AI model and trains the model. It retrieves the pre-processed and stored data from the database and supplies it to the generative AI model, updating the model with the latest information and training it. The input data is the data retrieved from the database, and the output data is the trained generative AI model.
[0703] Step 5:
[0704] The server maps multiple perspectives from the trained generative AI model into a unified space. Specifically, it aggregates data from different information sources and perspectives into a unified space using a clustering algorithm and visualizes it. The input data is the trained generative AI model, and the output data is the mapped perspective data.
[0705] Step 6:
[0706] The server receives requests from users and analyzes them. When a user sends a request from a device, the server analyzes the request using natural language processing (NLP). The input data is the user's request, and the output data is the analysis result.
[0707] Step 7:
[0708] The server searches and generates information based on the analyzed request, searches for relevant information from the database, and generates an appropriate answer using a generative AI model. The input data is the analysis result and the storage data saved in the database, and the output data is the generated answer.
[0709] Step 8:
[0710] The server recognizes the user's emotional state. It uses an emotion engine to analyze the user's emotional state (e.g., relief, anxiety, interest, etc.) from their search history and input text. The input data is the user's behavioral data, and the output data is the recognized emotional state.
[0711] Step 9:
[0712] The server adjusts the generated information based on the recognized emotional state, specifically by changing the tone of the text or filtering the information. The input data is the recognized emotional state and the generated response, and the output data is the adjusted response.
[0713] Step 10:
[0714] The terminal provides the generated information received from the server to the user. The terminal presents the information to the user in text format, notification format, etc. The input data is the adjusted response from the server, and the output data is the information received by the user.
[0715] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0716] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0717] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0718] [Third embodiment]
[0719] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0720] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0721] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0722] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0723] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0724] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0725] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0726] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0727] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0728] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0729] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0730] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0731] The present invention is a system that trains a generative AI model using data collected from public information sources and provides information from a unified perspective. This system consists of a server, a terminal, and a user.
[0732] Public information collection and preprocessing
[0733] First, the server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, it obtains the necessary data using APIs and web scraping technology. For example, it collects the latest policy information on COVID-19 from the government's open data portal and related articles from news sites. Next, the server preprocesses the collected data. It cleans the data, standardizes the format, removes special characters and spaces, and arranges the data into a unified format.
[0734] Storing in a database and training the generated AI model
[0735] The preprocessed data is stored in a database by the server. The server then retrieves the latest data from the database and inputs it into the generative AI model. This generative AI model has the ability to learn from the collected data and make inferences and predictions about new data. The server trains the generative AI model so that it always reflects the latest information.
[0736] Mapping to a unified space and calculating the relative viewpoints
[0737] The server runs a technology that maps multiple viewpoints into a unified space using a trained generative AI model. During this process, data from different sources and viewpoints is centrally aggregated. Furthermore, a clustering algorithm is used to calculate and visualize the relative positions of each viewpoint. As a result, users can easily grasp the relevance and reliability of the information.
[0738] Receiving user requests and generating information
[0739] A user requests specific information through a device (such as a PC or smartphone). For example, they send a request such as, "Please tell me the latest measures against COVID-19." The device then sends this request to the server. The server receives the request from the user and analyzes it using natural language processing (NLP). Based on the analyzed request, the server searches for relevant information from a unified space and generates appropriate comments or policy proposals using a generative AI model.
[0740] Providing information and concrete examples
[0741] The server sends the generated information to the device, which then displays the results to the user. For example, if a user requests, "I want to know about COVID-19 countermeasures," the server generates the latest policy information and suggestions, and sends information such as, "The latest government policy recommends vaccination in XX region. In addition, the home quarantine period has been changed from XX to XX days," to the device, which then displays this to the user. This system allows users to always easily access the latest, most reliable information, which can be used to understand policies and support their daily lives.
[0742] The processing flow will be explained below.
[0743] Step 1:
[0744] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.) by using REST APIs to retrieve data in JSON format or by using web scraping technology to obtain the required information.
[0745] Step 2:
[0746] The server pre-processes the data it receives, cleaning it (for example, removing whitespace and special characters), standardizing date and number formats, and arranging the data into a consistent format.
[0747] Step 3:
[0748] The server saves the pre-processed data to a database, where the information is stored efficiently and securely using INSERT operations into an SQL database.
[0749] Step 4:
[0750] The server retrieves the latest data from the database and inputs it into the generative AI model, which then uses the newly collected data to train the generative AI model.
[0751] Step 5:
[0752] The server trains the generative AI model, running machine learning algorithms that enable the model to make accurate predictions and inferences on new data.
[0753] Step 6:
[0754] The server maps multiple perspectives from the generative AI model into a unified space. Specifically, it plots data from different sources and perspectives in a multidimensional space and calculates their spatial relationships using a clustering algorithm.
[0755] Step 7:
[0756] The user requests specific information through the device, for example, "Please tell me the latest information on COVID-19 countermeasures," and sends it.
[0757] Step 8:
[0758] The terminal sends a request from the user to the server, and HTTP requests are generally used as the communication protocol.
[0759] Step 9:
[0760] The server analyzes the request received from the terminal, and uses natural language processing (NLP) to analyze the content of the request and identify its requirements.
[0761] Step 10:
[0762] The server searches for relevant information from the unified space based on the analysis results, leveraging the capabilities of SQL queries and generative AI models to extract the necessary information.
[0763] Step 11:
[0764] The server uses generative AI models to generate comments and policy proposals based on the retrieved information, using natural language generation (NLG) to provide the generated text in a format that is easy for users to understand.
[0765] Step 12:
[0766] The server then sends the generated information to the terminal. Again, HTTP responses are generally used as the communication protocol.
[0767] Step 13:
[0768] The device displays the information received from the server to the user, either by updating a web page or displaying the information in the mobile app UI, making it easily accessible to the user.
[0769] This series of steps allows users to efficiently utilize open public information and obtain the information they need quickly and accurately.
[0770] Example 1
[0771] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0772] In conventional information provision systems, when collecting data from public information sources and providing information from a unified perspective, multiple processes such as data preprocessing, training of generative AI models, and viewpoint mapping are required, and each process is often performed independently, resulting in problems that reduce the efficiency and accuracy of the entire system.In addition, there is a lack of functionality to properly calculate and visualize the positional relationships of viewpoints obtained from different information sources, making it difficult for users to quickly obtain reliable information.
[0773] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0774] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a data management device, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple viewpoints into a unified space from the trained generative AI model, means for calculating the positional relationship between viewpoints obtained from different information sources, means for receiving a request from a user, means for searching and generating information based on the request, and means for providing the generated information to the user, thereby enabling the user to quickly obtain reliable information and grasp a variety of information in a unified manner.
[0775] A "public information source" is a database or website that provides information that is publicly accessible on the Internet.
[0776] "Data preprocessing" is the process of cleaning, formatting, and removing unnecessary information from collected data to prepare it in a format that is easy to analyze and store.
[0777] A "data management device" refers to a database or storage system for efficiently and safely storing collected data.
[0778] A "generative artificial intelligence model" is a machine learning model that can learn from collected data and predict or generate data in the future.
[0779] "Training" is the learning process that a generative artificial intelligence model goes through to achieve optimal performance based on data.
[0780] A "unified space" is a data representation that centrally manages multiple perspectives collected from different sources and maps them in a form that makes comparison and analysis easy.
[0781] "Positional relationships between viewpoints" refers to information that indicates how multiple viewpoints are related to each other.
[0782] A "clustering algorithm" is a computational method for grouping data based on similarity.
[0783] "User Request" means a request for information entered by a user into the system.
[0784] "Information retrieval and generation means" refers to the process of locating appropriate information from databases and generative artificial intelligence models based on user requests and generating text, reports, etc. as needed.
[0785] "Generated Information" refers to information created by the system using a generative artificial intelligence model based on a user request.
[0786] The present invention is a system that trains a generative AI model using data collected from public information sources and provides information from a unified perspective. This system consists of a server, a terminal, and a user.
[0787] First, the server collects data from public sources on the internet, such as government databases, news sites, and social networking services. This process utilizes APIs and web scraping techniques, using tools such as Python's requests library and BeautifulSoup. For example, it obtains policy information about COVID-19 from government open data portals and collects related articles from news sites.
[0788] The server then preprocesses the collected data, which includes removing special characters and whitespace, cleaning, and formatting, using the Python pandas library to organize the data into a unified format.
[0789] The preprocessed data is stored in a data management device by the server. The server uses a database system such as MySQL or PostgreSQL to efficiently store the data. The stored data is then input into a generative artificial intelligence model to train the model. Machine learning frameworks such as TensorFlow and PyTorch are used for this training.
[0790] A trained generative artificial intelligence model maps multiple viewpoints into a unified space. The server integrates data from different sources and viewpoints and calculates the spatial relationships between viewpoints using a clustering algorithm in Python's Scikit-learn library. As a result, the relationships between viewpoints are visualized, and important information is centralized.
[0791] A user requests specific information through a device, such as a PC or smartphone. An example of a prompt would be "Please tell me the latest measures against COVID-19." The device then sends this request to the server, which analyzes it using natural language processing (NLP) tools. The server uses Python's spaCy or NLTK to analyze the request, searches for relevant information in a unified space based on the analysis results, and generates comments or policy proposals using a generative artificial intelligence model if necessary.
[0792] Finally, the server sends the generated information to the device, which then displays the results to the user. For example, information such as "The latest government policy recommends vaccination in the ____ region. Furthermore, the self-isolation period has been changed from ____ to ____ days" can be provided. In this way, users can quickly obtain the latest and most reliable information.
[0793] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0794] Step 1: Gather public information
[0795] The server collects data from public sources on the Internet, including government databases, news sites, social networking services, etc. The server uses Python's requests library to retrieve the required data from APIs, and also uses BeautifulSoup to scrape articles from news sites.
[0796] Input: URL or API endpoint of a public source
[0797] Output: Raw data (text, JSON, etc.)
[0798] Step 2: Preprocessing the data
[0799] The server preprocesses the collected data, which includes removing unnecessary whitespace and special characters, standardizing the format, etc. The server uses the Python pandas library to clean the data and convert it into a unified format.
[0800] Input: Raw data
[0801] Output: Preprocessed data (unified format)
[0802] Step 3: Saving to the database
[0803] The server stores the preprocessed data in a data management device (database), such as MySQL or PostgreSQL. The server establishes a connection to the database and inserts data using SQL queries.
[0804] Input: Preprocessed data
[0805] Output: Data stored in the database
[0806] Step 4: Training the generative AI model
[0807] The server retrieves the latest data from the database and trains a generative AI model using machine learning frameworks such as TensorFlow and PyTorch. The server inputs the data into the model and trains it over multiple epochs.
[0808] Input: Data retrieved from the database
[0809] Output: A trained generative AI model
[0810] Step 5: Mapping to a unified space
[0811] The server uses a trained generative AI model to map multiple viewpoints into a unified space using a clustering algorithm using Python's Scikit-learn library, which calculates the spatial relationships between viewpoints obtained from different sources.
[0812] Input: A trained generative AI model and associated data
[0813] Output: viewpoint data mapped to a unified space
[0814] Step 6: Calculate the viewpoint position
[0815] The server calculates the relative positions of the viewpoints mapped to the unified space and prepares for visualization. It calculates the distance and relationship between viewpoints and generates a network graph. This process uses the Python networkx library.
[0816] Input: Mapped viewpoint data
[0817] Output: Data for visualizing the relative positions of viewpoints
[0818] Step 7: Receiving a user request
[0819] The user requests specific information through the device, for example by entering a prompt such as "Please tell me the latest information on COVID-19 countermeasures." The device then sends this request to the server.
[0820] Input: User request (prompt sentence)
[0821] Output: Request sent to the server
[0822] Step 8: Generate information
[0823] The server analyzes requests received from users using natural language processing (NLP) tools, such as Python's spaCy and NLTK. Based on the analysis results, the server searches for relevant information in a unified space and generates appropriate comments and policy proposals using a generative AI model.
[0824] Input: The request sent to the server
[0825] Output: Generated information (comments and policy suggestions)
[0826] Step 9: Provide information
[0827] The server sends the generated information to the device, which then displays the results to the user, such as, "The latest government policy recommends vaccination in the ____ region. In addition, the home quarantine period has been changed from ____ to ____ days."
[0828] Input: Generated information
[0829] Output: Information displayed on the user's terminal
[0830] In this way, specific actions are performed at each step, allowing the user to quickly obtain reliable, up-to-date information.
[0831] (Application example 1)
[0832] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0833] In today's world, businesses and individuals are constantly exposed to the threat of cyberattacks, resulting in increased risk of serious economic losses and information leaks. Cyberthreats evolve daily, and information about them is provided fragmentedly from numerous sources, making it difficult to unify, understand, and respond immediately. Furthermore, the lack of a system that can assess risks in real time and issue immediate alerts for high-risk threats means that prompt responses are delayed, potentially resulting in greater damage. To address these challenges and improve the effectiveness of cybersecurity, a system is needed that can unify the latest information, quickly and accurately assess risks, and notify users in real time.
[0834] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0835] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a database, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple perspectives into a unified space from the trained generative AI model, means for receiving requests from users, means for searching and generating information based on the requests, means for providing the generated information to users, means for risk assessing cybersecurity-related information, and means for issuing alerts to users about high-risk threats in real time. This enables rapid and accurate risk assessment of cyberattacks and the issuance of alerts in real time, thereby strengthening the security measures of companies and individuals.
[0836] "Public sources" refer to data sources that are accessible to anyone and from which information can be obtained, such as government databases, news sites, and social media.
[0837] "Data preprocessing" is the process of cleaning collected data, standardizing the format, removing unnecessary special characters, etc., to improve the quality of the data.
[0838] "Database" refers to a collection of data that stores collected and preprocessed data and structures it for easy access and retrieval.
[0839] A "generative AI model" is an artificial intelligence model that learns from collected data and has the ability to make inferences and predictions about new data.
[0840] A "unified space" refers to a data space that centrally aggregates data obtained from different sources and perspectives and presents it in a unified format and perspective.
[0841] A "clustering algorithm" is an algorithm used to group data based on similarity and calculate the relative positions of viewpoints.
[0842] "Request" means a request or inquiry sent by a User to a Server for specific information.
[0843] "Risk assessment" is the process of using collected data and trained generative AI models to assess the risk of potential cybersecurity-related threats and attacks.
[0844] "Real-time alerts" refer to warning messages that immediately notify users when cybersecurity threats increase and encourage them to take prompt action.
[0845] This invention is a system that trains a generative AI model using data collected from public sources and provides information from a unified perspective. The system functions through the following stages: [collection of public information], [data preprocessing], [storage in a database], [training of a generative AI model], [mapping to a unified space], [issuance of risk assessment and real-time alerts], and [provision of a user interface].
[0846] Hardware and Software Configuration
[0847] Hardware:
[0848] Server: Responsible for data collection, storage, model training, inference, and request processing.
[0849] Devices: Smartphones, smart glasses, head-mounted displays, robots, etc. Serve as the interface with the user.
[0850] software:
[0851] Data collection: API, web scraping tools (BeautifulSoup, Selenium)
[0852] Data preprocessing: Python (Pandas, NumPy)
[0853] Model training: TensorFlow, PyTorch
[0854] Clustering: Scikit-learn
[0855] Natural Language Processing: SpaCy, NLTK
[0856] User interface: Flutter (smartphones), Unity (smart glasses, head-mounted displays, robots)
[0857] Data collection and preprocessing
[0858] The server uses APIs and web scraping tools to collect data from public sources on the Internet (government databases, news sites, social media, etc.). The collected data is pre-processed using Python libraries to clean and standardize the format. The pre-processed data is then structured and stored in a database.
[0859] Generative AI models and mapping to a unified space
[0860] The saved data is input into a generative AI model, which is then trained. The generative AI model is developed using TensorFlow and PyTorch. The trained generative AI model maps the data obtained from multiple viewpoints into a unified space. The relative positions of the viewpoints are calculated and visualized using Scikit-learn's clustering algorithm.
[0861] Risk Assessment and Real-Time Alerts
[0862] The server evaluates cybersecurity-related risks based on the generated information and sends real-time alerts to users' systems about high-risk threats, allowing them to take prompt action.
[0863] User Interface
[0864] User requests are sent to the server via the device. The request is analyzed and relevant information is searched for and generated. User interfaces developed using Flutter and Unity are provided to users via smartphones, smart glasses, head-mounted displays, and robots.
[0865] Examples of concrete examples and prompts
[0866] For example, if a new ransomware attack is reported, the server collects relevant information, which is then scrutinized by the generative AI model and sent to the device, where it is notified to the user visually and audibly, providing immediate countermeasures.
[0867] Prompt Sentence Examples
[0868] User: "What's the latest ransomware attack?"
[0869] Security Assistant AI: "The latest ransomware attack is called ____ and occurred on ____ / ____ / __. Details of the attack are as follows: ____. As a countermeasure, we recommend ____."
[0870] This allows users to respond quickly to the latest cyber threats and strengthen their security.
[0871] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0872] Step 1:
[0873] The server uses APIs and web scraping tools (BeautifulSoup, Selenium, etc.) to collect data from public sources (government databases, news sites, social media, etc.). The input of this step is the public source, and the output is the collected raw data. Specifically, the server accesses the specified URL and retrieves the required information.
[0874] Step 2:
[0875] The server performs preprocessing on the collected raw data. Preprocessing includes data cleaning (removing unnecessary data and special characters), formatting unification (e.g., changing the date and time format), and text standardization (e.g., lowercasing). The input of this step is the collected raw data, and the output is preprocessed clean data. Specifically, the data is converted using Python's Pandas and NumPy libraries.
[0876] Step 3:
[0877] The server saves the preprocessed data in a database. The input of this step is clean data, and the output is data stored in the database. Specifically, it inserts and updates data in the database using SQL queries.
[0878] Step 4:
[0879] The server inputs the clean data into the generative AI model and trains the model. The input for this step is clean data retrieved from the database, and the output is a trained generative AI model. Specifically, the generative AI model is trained using TensorFlow or PyTorch.
[0880] Step 5:
[0881] The server uses a trained generative AI model to map multiple viewpoints into a unified space and calculates the relative positions of the viewpoints using Scikit-learn's clustering algorithm. The input to this step is the generative AI model and clean data, and the output is viewpoint data mapped into a unified space. Specifically, the clustering algorithm is applied to group the data.
[0882] Step 6:
[0883] The user sends an information request to the server through their device. The input in this step is the request from the user, and the output is the analysis result of the request content. Specifically, the user inputs the request using a smartphone or smart glasses.
[0884] Step 7:
[0885] Based on the received request, the server analyzes the request content using natural language processing (SpaCy or NLTK) and searches for related information from a unified space. The input to this step is the analysis result, and the output is the searched related information. Specifically, it extracts related information from a database based on the request content.
[0886] Step 8:
[0887] The server generates information using a generative AI model based on the searched related information and sends it to the terminal. The input of this step is the searched related information, and the output is the generated information. Specifically, the server generates information based on the generative AI model and sends it to the terminal.
[0888] Step 9:
[0889] The terminal provides the generated information received from the server to the user. The input of this step is the generated information from the server, and the output is information displayed and notified in a format that the user can understand. Specific operations include providing information to the user visually or audibly using a smartphone, smart glasses, a head-mounted display, or a robot.
[0890] Through this series of processing steps, the system alerts users in real time to high-risk cyber threats, enabling them to respond quickly.
[0891] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0892] This invention is a system that trains a generative AI model using data collected from public sources and provides information from a unified perspective. This system also incorporates an emotion engine that recognizes user emotions, improving the effectiveness of information provision. The system consists of a server, a terminal, and a user.
[0893] Public information collection and preprocessing
[0894] First, the server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, it obtains the necessary data using APIs and web scraping technology. For example, it collects the latest policy information on COVID-19 from the government's open data portal and related articles from news sites. Next, the server preprocesses the collected data. It cleans the data, standardizes the format, removes special characters and spaces, and arranges the data into a unified format.
[0895] Storing in a database and training the generated AI model
[0896] The preprocessed data is stored in a database by the server. The server then retrieves the latest data from the database and inputs it into the generative AI model. This generative AI model has the ability to learn from the collected data and make inferences and predictions about new data. The server trains the generative AI model so that it always reflects the latest information.
[0897] Mapping to a unified space and calculating the relative viewpoints
[0898] The server runs a technology that maps multiple viewpoints into a unified space using a trained generative AI model. During this process, data from different sources and viewpoints is centrally aggregated. Furthermore, a clustering algorithm is used to calculate and visualize the relative positions of each viewpoint. As a result, users can easily grasp the relevance and reliability of the information.
[0899] Emotion engine recognizes user emotions
[0900] The device sends user input and behavioral data (e.g., search history, clicks, input text, etc.) to the emotion engine, which analyzes the data and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.). This analysis is performed using natural language processing (NLP) and machine learning algorithms.
[0901] Receiving user requests and generating information
[0902] A user requests specific information through a device (such as a PC or smartphone). For example, they send a request such as, "Please tell me the latest measures against COVID-19." The device then sends this request to the server. The server receives the request from the user and analyzes it using natural language processing (NLP). Based on the analyzed request, the server searches for relevant information from a unified space and generates appropriate comments or policy proposals using a generative AI model.
[0903] Customize information based on user sentiment
[0904] Based on the analysis results from the emotion engine, the server adjusts the tone and content of the information generated by the generative AI model. For example, if the user is feeling anxious, the information provided will be generated in a reassuring tone. In this way, information is customized according to the user's emotional state, making it possible to provide more effective information.
[0905] Providing information and concrete examples
[0906] The server sends the generated information to the device, which then displays the results to the user. For example, if a user requests, "I want to know about COVID-19 countermeasures," the server generates the latest policy information and suggestions, and based on the results of the emotion engine, sends information such as, "The latest government policy recommends vaccination in the XX region. In addition, the self-isolation period has been changed from XX to XX days. If you are concerned, please refer to the detailed guidelines." to the device, which then displays this to the user. This system allows users to easily access the latest and most reliable information, which can be used to understand policies and support their daily lives. Furthermore, providing information based on the user's emotions increases user satisfaction and trust.
[0907] The processing flow will be explained below.
[0908] Step 1:
[0909] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.) by using REST APIs to retrieve data in JSON format or by using web scraping technology to obtain the required information.
[0910] Step 2:
[0911] The server pre-processes the data it receives, cleaning it (for example, removing whitespace and special characters), standardizing date and number formats, and arranging the data into a consistent format.
[0912] Step 3:
[0913] The server saves the pre-processed data to a database, where the information is stored efficiently and securely using INSERT operations into an SQL database.
[0914] Step 4:
[0915] The server retrieves the latest data from the database and inputs it into the generative AI model, which then uses the newly collected data to train the generative AI model.
[0916] Step 5:
[0917] The server trains the generative AI model, running machine learning algorithms that enable the model to make accurate predictions and inferences on new data.
[0918] Step 6:
[0919] The server maps multiple perspectives from the generative AI model into a unified space. Specifically, it plots data from different sources and perspectives in a multidimensional space and calculates their spatial relationships using a clustering algorithm.
[0920] Step 7:
[0921] The terminal sends the user's input data and behavioral data (e.g., past search history, click status, entered text, etc.) to the emotion engine.
[0922] Step 8:
[0923] The emotion engine analyzes data sent from the device and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.) using natural language processing (NLP) and machine learning algorithms.
[0924] Step 9:
[0925] Users request specific information through their device, for example, by typing in a request such as "Please tell me the latest information on COVID-19 countermeasures" and submitting it.
[0926] Step 10:
[0927] The terminal sends a request from the user to the server, and HTTP requests are generally used as the communication protocol.
[0928] Step 11:
[0929] The server analyzes the request received from the terminal, and uses natural language processing (NLP) to analyze the content of the request and identify its requirements.
[0930] Step 12:
[0931] The server searches for relevant information from the unified space based on the analysis results, leveraging the capabilities of SQL queries and generative AI models to extract the necessary information.
[0932] Step 13:
[0933] The server uses generative AI models to generate comments and policy proposals based on the retrieved information, using natural language generation (NLG) to provide the generated text in a format that is easy for users to understand.
[0934] Step 14:
[0935] The server adjusts the tone and content of the generated information based on the user's emotion analysis results obtained from the emotion engine. For example, if the user is feeling anxious, the information provided will be generated in a reassuring tone.
[0936] Step 15:
[0937] The server then sends the generated information to the terminal. Again, HTTP responses are generally used as the communication protocol.
[0938] Step 16:
[0939] The device displays the information received from the server to the user, either by updating a web page or displaying the information in the mobile app UI, making it easily accessible to the user.
[0940] This series of steps allows users to efficiently utilize public information and obtain the information they need quickly and accurately, as well as receive information tailored to their own emotions.
[0941] Example 2
[0942] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0943] In conventional information provision systems, it has been difficult to collect data from multiple information sources and provide users with reliable information from a unified perspective. Furthermore, the information provided does not take into account the user's emotional state, resulting in a decrease in user satisfaction. The present invention aims to solve these problems by handling data collected from a variety of information sources in a unified manner and providing information that corresponds to the user's emotional state.
[0944] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0945] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a database, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple perspectives into a unified space from the trained generative AI model, means for analyzing emotions using user input data and behavioral data, means for receiving requests from users, means for searching and generating information based on the requests, means for providing the generated information to the user, and means for adjusting the tone and content of the generated information based on the results of the emotion analysis. This makes it possible to handle data collected from multiple information sources in a unified manner and provide appropriate information according to the user's emotional state.
[0946] "Public information sources" refer to information sources such as databases, news sites, and social media that are publicly available on the Internet.
[0947] "Means of collecting data" refers to the processes and techniques used to obtain the required information from public sources using APIs and web scraping techniques.
[0948] "Preprocessing" refers to the process of removing noise from collected data, standardizing the format, removing special characters and unnecessary spaces, and otherwise preparing the data for easier analysis.
[0949] "Means of database storage" refers to the process or technique by which pre-processed data is stored in a data management system.
[0950] A "generative AI model" refers to an artificial intelligence model that has the ability to learn from large amounts of data and make inferences and predictions based on new data.
[0951] "Training means" refers to the process or technique of inputting data into a generative AI model and optimizing the model's parameters.
[0952] "Means of mapping to a unified space" refers to the process of centrally aggregating data from multiple perspectives obtained by a generative AI model and visualizing the relationships between the perspectives in a low-dimensional space.
[0953] "Means of emotion analysis" refers to natural language processing and machine learning technologies that recognize users' emotional states based on their input data and behavioral data.
[0954] "Means for receiving requests from users" refers to the process by which the server receives inquiries or information requests sent by users through their terminals.
[0955] "Means for information retrieval and generation" refers to the process of analyzing a user's request, searching for relevant information from a database based on that request, and providing appropriate information using a generative AI model.
[0956] "Means for providing information to a user" refers to the process of displaying the generated information to a user through a terminal.
[0957] "Measures to adjust tone and content" refers to the process of changing the expression and content of generated information based on the results of sentiment analysis in accordance with the user's emotional state.
[0958] This system trains a generative AI model using data collected from public sources to provide information from a unified perspective. It also incorporates an emotion engine that recognizes the user's emotional state and provides information accordingly. The system consists of a server, a terminal, and a user.
[0959] Public information collection and preprocessing
[0960] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, the server obtains data using APIs and web scraping technologies (e.g., BeautifulSoup or Scrapy). For example, the server collects the latest policy information about COVID-19 from the government's open data portal and related articles from news sites. The server then preprocesses the collected data. This involves removing noise, standardizing the format, and removing special characters and spaces, thereby arranging the data into a unified format. This makes it easier to input into the generative AI model.
[0961] Storing in a database and training the generated AI model
[0962] The preprocessed data is stored in a database by the server. Specifically, the server connects to a database system such as MySQL or MongoDB and stores the data. The server then retrieves the latest data from the database and inputs it into a generative AI model (e.g., GPT-3 or BERT). This generative AI model has the ability to learn from the collected data and make inferences and predictions based on new data. The server trains the generative AI model using libraries such as TensorFlow and PyTorch to ensure that it always reflects the latest information.
[0963] Mapping information to a unified space and calculating spatial relationships
[0964] The server maps multiple perspectives obtained from the trained generative AI model into a unified space. In this process, data obtained from different information sources and perspectives is centrally aggregated. The server then uses a clustering algorithm (e.g., k-means) to calculate the relative positions of each perspective and visualize them in a low-dimensional space. This allows users to easily grasp the relevance and reliability of the information.
[0965] Emotion engine recognizes user emotions
[0966] The device sends the user's input data and behavioral data (e.g., search history, clicks, input text, etc.) to the emotion engine. The emotion engine analyzes this data and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.). Specifically, it performs the analysis using natural language processing (NLP) and machine learning algorithms (e.g., the emotion analysis model VADER).
[0967] Receiving user requests and generating information
[0968] A user requests specific information through a device (e.g., a PC or smartphone). For example, the user types, "Please tell me the latest measures against COVID-19." This request is sent from the device to the server. The server receives the request and analyzes it using natural language processing (NLP) technology. Based on the analysis results, the server searches for relevant information from a unified space and generates appropriate information using a generative AI model.
[0969] Regulating and providing information based on emotions
[0970] Based on the analysis results from the emotion engine, the server adjusts the tone and content of the generated information. For example, if the user is feeling anxious, the server will provide information in a reassuring tone, such as, "Don't worry, the latest government policy recommends vaccination in the XX area." The final information is sent from the server to the device, which then displays it to the user. This allows users to always have easy access to the latest and most reliable information, allowing them to use the information with greater peace of mind.
[0971] Examples of concrete examples and prompts
[0972] For example, if you receive a request such as "I want to know the latest information on COVID-19 countermeasures," the prompt might look like this:
[0973] What is the latest information on COVID-19 prevention measures? My current emotional state is anxiety.
[0974] By inputting this prompt into the generative AI model, appropriate information is provided.
[0975] This system provides more effective information provision by customizing information according to the user's emotional state, and by centralizing data collected from various sources, users can quickly obtain reliable information.
[0976] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0977] Step 1:
[0978] The server collects data from public sources (government databases, news sites, social media, etc.). Specifically, the server obtains information using APIs or web scraping technologies (e.g., BeautifulSoup or Scrapy). The input is the URL or API key of the public source, and the output is the collected raw data. For example, the server obtains the latest policy information about COVID-19 from the government's open data portal.
[0979] Step 2:
[0980] The server preprocesses the collected data. The input is the raw data collected in step 1, and the output is clean, uniformly formatted data. Specifically, the server uses regular expressions to remove noise, standardize the format, and remove special characters and whitespace. This prepares the data in a form that is easy to input into the generative AI model.
[0981] Step 3:
[0982] The server saves the preprocessed data in a database. The input is the preprocessed data, and the output is the data stored in the database. Specifically, the server connects to a database system such as MySQL or MongoDB and executes SQL statements to insert data. This allows the data to be managed in an organized manner.
[0983] Step 4:
[0984] The server retrieves the latest data from the database and inputs it into the generative AI model. The input is the formatted data retrieved from the database, and the output is the data processed by the generative AI model. Specifically, the server retrieves the data and inputs it into the generative AI model (e.g., GPT-3 or BERT) using libraries such as TensorFlow or PyTorch.
[0985] Step 5:
[0986] The server trains the generative AI model. The input is the training data and the model's initial parameters, and the output is the trained model. For example, the server uses a dataset to optimize the model's parameters and improve its inference ability on new data. This training process is performed using libraries such as TensorFlow and PyTorch.
[0987] Step 6:
[0988] The server maps multiple viewpoints from a trained generative AI model into a unified space. The input is viewpoint data generated by the model, and the output is a unified viewpoint space. Specifically, the server visualizes the viewpoint data in a low-dimensional space using UMAP (Uniform Manifold Approximation and Projection) or a clustering algorithm (e.g., k-means).
[0989] Step 7:
[0990] The device sends the user's input data and behavioral data to the emotion engine. The input is the user's input data (search history, clicks, input text, etc.), and the output is the emotion analysis results. The device sends this data to the emotion engine, which then analyzes it using natural language processing technology.
[0991] Step 8:
[0992] A user requests specific information through a device. The input is the user's request (e.g., "Please tell me the latest measures against COVID-19"), and the output is the request data. This request is sent from the device to the server.
[0993] Step 9:
[0994] The server receives the user's request and analyzes it using natural language processing technology. The input is the text data of the user's request, and the output is the analysis result. Based on the analysis result, the server searches for relevant information from the unified space and generates appropriate information using a generative AI model.
[0995] Step 10:
[0996] The server adjusts the tone and content of the generated information based on the analysis results from the emotion engine. The input is the emotion analysis result and the generated information, and the output is the adjusted information. For example, if the user is feeling anxious, the server will provide information in a reassuring tone.
[0997] Step 11:
[0998] The server sends the generated information to the terminal, which then displays it to the user. The input is the adjusted information, and the output is the information displayed to the user. The terminal displays on the user's smartphone or PC, "The latest government policy recommends vaccination in the XX area. Additionally, the home quarantine period has been changed from XX to XX days."
[0999] (Application example 2)
[1000] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1001] In modern society, the importance of security information is increasing day by day. However, there is still a lack of systems that allow users to obtain reliable security information in real time and provide that information tailored to the user's emotional state. Furthermore, there is the challenge of properly recognizing users' emotions, such as anxiety and fear, and providing customized information based on those emotions. Therefore, it is necessary to provide a system that allows users to receive security information with peace of mind.
[1002] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1003] In this invention, the server includes means for collecting data from public information sources, means for pre-processing the collected data, and means for storing the pre-processed data in a database, thereby enabling security information collected from public information sources to be provided to the user in real time and further customized by recognizing the user's emotional state.
[1004] "Public sources" are sources that are free or publicly available on the internet and can be accessed by anyone. Examples include government databases, news sites, and social media.
[1005] "Methods of data collection" refers to technologies and tools used to automatically obtain the required data from public sources on the Internet, including the use of APIs and web scraping techniques.
[1006] "Preprocessing methods" refers to techniques for improving the quality of collected data and arranging it into a unified format, such as data cleaning, formatting standardization, and removal of special characters.
[1007] "Means of storing data in a database" refers to a system that efficiently stores pre-processed data and manages it for easy retrieval and use. Examples include RDBMS and NoSQL databases.
[1008] "Means of input to generative AI models" refers to technologies that provide data stored in a database to generative AI models for use in training and prediction.
[1009] "Training methods" refers to the techniques used to train generative AI models with the latest data and improve their performance, including machine learning algorithms and data feeds.
[1010] "Means of mapping to a unified space" refers to techniques that unify data from different sources and perspectives to allow users to understand the relevance and reliability of the information. Specifically, this includes dimensionality reduction and clustering algorithms.
[1011] "Means for receiving requests" refers to the technology used to obtain information requests from users and communicate them to the system, including the user interface and the NLP engine.
[1012] "Information search and generation means" refers to technologies that search for relevant information from databases based on user requests and generate appropriate answers or comments, including search algorithms and generative AI models.
[1013] "Means for providing to users" refers to technologies for presenting generated information to users in an easy-to-understand manner, including user interfaces and notification systems.
[1014] "Emotional state recognition" refers to technologies for identifying a user's current emotional state based on their input and behavioral data, including natural language processing and machine learning algorithms.
[1015] "Information tailoring measures" refers to techniques for adjusting the tone and content of information provided in response to the perceived emotional state of the user, including tone modification and information filtering.
[1016] This invention is a system that collects data from public sources, uses a generative AI model to generate information requested by users, and further customizes the information according to the user's emotional state. The system mainly consists of three elements: a server, a terminal, and a user.
[1017] System configuration
[1018] server
[1019] The server has the following main functions:
[1020] Data collection from public sources: The server collects data in real time from public sources on the Internet, using APIs and web scraping techniques to extract necessary security information from government databases, news sites, social media, etc.
[1021] Data preprocessing: The collected data is preprocessed by data cleaning, standardizing the format, removing special characters, etc.
[1022] Database storage: The preprocessed data is stored in a database that can be efficiently searched and updated, typically using an RDBMS or NoSQL database.
[1023] Information generation using a generative AI model: The stored data is input into a generative AI model to train the model. The trained generative AI model is then used to generate information based on user requests.
[1024] Emotion recognition: Analyzes user input data (e.g., search history, clicks, input text) and uses an emotion engine to recognize the user's current emotional state.
[1025] Information customization: Adjusting the tone and content of generated information based on perceived emotional state.
[1026] Terminal
[1027] The device (e.g., smartphone or PC) has the following features:
[1028] Sending a user request: A user requests specific information and sends it to the server via their device. For example, a request could be, "Please tell me about the latest cyber-attack information."
[1029] Receiving and displaying information: Receives information generated by the server and displays it to the user. The display method can be in various formats such as text or notification.
[1030] user
[1031] The user does the following:
[1032] Information Request: Using a smartphone or PC, a user requests specific information. This request is made in natural language and sent to the server.
[1033] Review the information: Review the information received and take further action as necessary.
[1034] Specific examples of processing
[1035] When a user requests, "I'd like to know about recent cyber-attack trends," the server first collects the latest security-related data from public sources. This data is preprocessed and stored in a database. Then, a generative AI model is used to generate appropriate information and customize it based on the user's emotional state. Finally, the generated, specific answer (e.g., "Recent cyber-attacks have been characterized by large-scale phishing campaigns and the emergence of new ransomware. In particular, there has been an increase in techniques that identify and encrypt sensitive data within companies and demand ransom. Effective defensive measures include regular software updates and strengthened employee education. If you have concerns, please refer to our more detailed security policy") is provided to the user via their device.
[1036] Prompt Sentence Examples
[1037] Examples of specific prompts based on user input include:
[1038] User input: "I'd like to know about recent trends in cyber attacks."
[1039] Prompt: "User's feelings: anxious. Collect the security-related information: (Recently collected security information). User's request: I would like to know about recent cyber-attack trends."
[1040] This specific example allows users to obtain the latest and most reliable security information in real time, while also providing specific measures to alleviate their concerns.
[1041] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1042] Step 1:
[1043] The server collects data from public sources. Specifically, it uses APIs and web scraping technology to obtain real-time data from government databases, news sites, and social media on the Internet. The collected data is temporarily stored as raw data. Inputs include the URL and API key of the target public source, and the obtained raw data is included as output.
[1044] Step 2:
[1045] The server preprocesses the collected data. Specific operations include data cleaning (removing missing values and outliers), standardizing formats (e.g., standardizing character codes and date formats), removing special characters, etc. The input data is the raw data collected in the previous step, and the output data is the preprocessed clean data.
[1046] Step 3:
[1047] The server stores the preprocessed data in a database. Specifically, it uses an RDBMS or NoSQL database to efficiently store data. The input data is the preprocessed clean data, and the output is the storage data stored in the database.
[1048] Step 4:
[1049] The server inputs the stored data into the generative AI model and trains the model. It retrieves the pre-processed and stored data from the database and supplies it to the generative AI model, updating the model with the latest information and training it. The input data is the data retrieved from the database, and the output data is the trained generative AI model.
[1050] Step 5:
[1051] The server maps multiple perspectives from the trained generative AI model into a unified space. Specifically, it aggregates data from different information sources and perspectives into a unified space using a clustering algorithm and visualizes it. The input data is the trained generative AI model, and the output data is the mapped perspective data.
[1052] Step 6:
[1053] The server receives requests from users and analyzes them. When a user sends a request from a device, the server analyzes the request using natural language processing (NLP). The input data is the user's request, and the output data is the analysis result.
[1054] Step 7:
[1055] The server searches and generates information based on the analyzed request, searches for relevant information from the database, and generates an appropriate answer using a generative AI model. The input data is the analysis result and the storage data saved in the database, and the output data is the generated answer.
[1056] Step 8:
[1057] The server recognizes the user's emotional state. It uses an emotion engine to analyze the user's emotional state (e.g., relief, anxiety, interest, etc.) from their search history and input text. The input data is the user's behavioral data, and the output data is the recognized emotional state.
[1058] Step 9:
[1059] The server adjusts the generated information based on the recognized emotional state, specifically by changing the tone of the text or filtering the information. The input data is the recognized emotional state and the generated response, and the output data is the adjusted response.
[1060] Step 10:
[1061] The terminal provides the generated information received from the server to the user. The terminal presents the information to the user in text format, notification format, etc. The input data is the adjusted response from the server, and the output data is the information received by the user.
[1062] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1063] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1064] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1065] [Fourth embodiment]
[1066] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1067] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1068] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1069] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1070] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1071] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1072] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1073] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1074] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1075] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1076] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1077] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1078] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1079] The present invention is a system that trains a generative AI model using data collected from public information sources and provides information from a unified perspective. This system consists of a server, a terminal, and a user.
[1080] Public information collection and preprocessing
[1081] First, the server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, it obtains the necessary data using APIs and web scraping technology. For example, it collects the latest policy information on COVID-19 from the government's open data portal and related articles from news sites. Next, the server preprocesses the collected data. It cleans the data, standardizes the format, removes special characters and spaces, and arranges the data into a unified format.
[1082] Storing in a database and training the generated AI model
[1083] The preprocessed data is stored in a database by the server. The server then retrieves the latest data from the database and inputs it into the generative AI model. This generative AI model has the ability to learn from the collected data and make inferences and predictions about new data. The server trains the generative AI model so that it always reflects the latest information.
[1084] Mapping to a unified space and calculating the relative viewpoints
[1085] The server runs a technology that maps multiple viewpoints into a unified space using a trained generative AI model. During this process, data from different sources and viewpoints is centrally aggregated. Furthermore, a clustering algorithm is used to calculate and visualize the relative positions of each viewpoint. As a result, users can easily grasp the relevance and reliability of the information.
[1086] Receiving user requests and generating information
[1087] A user requests specific information through a device (such as a PC or smartphone). For example, they send a request such as, "Please tell me the latest measures against COVID-19." The device then sends this request to the server. The server receives the request from the user and analyzes it using natural language processing (NLP). Based on the analyzed request, the server searches for relevant information from a unified space and generates appropriate comments or policy proposals using a generative AI model.
[1088] Providing information and concrete examples
[1089] The server sends the generated information to the device, which then displays the results to the user. For example, if a user requests, "I want to know about COVID-19 countermeasures," the server generates the latest policy information and suggestions, and sends information such as, "The latest government policy recommends vaccination in XX region. In addition, the home quarantine period has been changed from XX to XX days," to the device, which then displays this to the user. This system allows users to always easily access the latest, most reliable information, which can be used to understand policies and support their daily lives.
[1090] The processing flow will be explained below.
[1091] Step 1:
[1092] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.) by using REST APIs to retrieve data in JSON format or by using web scraping technology to obtain the required information.
[1093] Step 2:
[1094] The server pre-processes the data it receives, cleaning it (for example, removing whitespace and special characters), standardizing date and number formats, and arranging the data into a consistent format.
[1095] Step 3:
[1096] The server saves the pre-processed data to a database, where the information is stored efficiently and securely using INSERT operations into an SQL database.
[1097] Step 4:
[1098] The server retrieves the latest data from the database and inputs it into the generative AI model, which then uses the newly collected data to train the generative AI model.
[1099] Step 5:
[1100] The server trains the generative AI model, running machine learning algorithms that enable the model to make accurate predictions and inferences on new data.
[1101] Step 6:
[1102] The server maps multiple perspectives from the generative AI model into a unified space. Specifically, it plots data from different sources and perspectives in a multidimensional space and calculates their spatial relationships using a clustering algorithm.
[1103] Step 7:
[1104] The user requests specific information through the device, for example, "Please tell me the latest information on COVID-19 countermeasures," and sends it.
[1105] Step 8:
[1106] The terminal sends a request from the user to the server, and HTTP requests are generally used as the communication protocol.
[1107] Step 9:
[1108] The server analyzes the request received from the terminal, and uses natural language processing (NLP) to analyze the content of the request and identify its requirements.
[1109] Step 10:
[1110] The server searches for relevant information from the unified space based on the analysis results, leveraging the capabilities of SQL queries and generative AI models to extract the necessary information.
[1111] Step 11:
[1112] The server uses generative AI models to generate comments and policy proposals based on the retrieved information, using natural language generation (NLG) to provide the generated text in a format that is easy for users to understand.
[1113] Step 12:
[1114] The server then sends the generated information to the terminal. Again, HTTP responses are generally used as the communication protocol.
[1115] Step 13:
[1116] The device displays the information received from the server to the user, either by updating a web page or displaying the information in the mobile app UI, making it easily accessible to the user.
[1117] This series of steps allows users to efficiently utilize open public information and obtain the information they need quickly and accurately.
[1118] Example 1
[1119] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1120] In conventional information provision systems, when collecting data from public information sources and providing information from a unified perspective, multiple processes such as data preprocessing, training of generative AI models, and viewpoint mapping are required, and each process is often performed independently, resulting in problems that reduce the efficiency and accuracy of the entire system.In addition, there is a lack of functionality to properly calculate and visualize the positional relationships of viewpoints obtained from different information sources, making it difficult for users to quickly obtain reliable information.
[1121] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1122] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a data management device, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple viewpoints into a unified space from the trained generative AI model, means for calculating the positional relationship between viewpoints obtained from different information sources, means for receiving a request from a user, means for searching and generating information based on the request, and means for providing the generated information to the user, thereby enabling the user to quickly obtain reliable information and grasp a variety of information in a unified manner.
[1123] A "public information source" is a database or website that provides information that is publicly accessible on the Internet.
[1124] "Data preprocessing" is the process of cleaning, formatting, and removing unnecessary information from collected data to prepare it in a format that is easy to analyze and store.
[1125] A "data management device" refers to a database or storage system for efficiently and safely storing collected data.
[1126] A "generative artificial intelligence model" is a machine learning model that can learn from collected data and predict or generate data in the future.
[1127] "Training" is the learning process that a generative artificial intelligence model goes through to achieve optimal performance based on data.
[1128] A "unified space" is a data representation that centrally manages multiple perspectives collected from different sources and maps them in a form that makes comparison and analysis easy.
[1129] "Positional relationships between viewpoints" refers to information that indicates how multiple viewpoints are related to each other.
[1130] A "clustering algorithm" is a computational method for grouping data based on similarity.
[1131] "User Request" means a request for information entered by a user into the system.
[1132] "Information retrieval and generation means" refers to the process of locating appropriate information from databases and generative artificial intelligence models based on user requests and generating text, reports, etc. as needed.
[1133] "Generated Information" refers to information created by the system using a generative artificial intelligence model based on a user request.
[1134] The present invention is a system that trains a generative AI model using data collected from public information sources and provides information from a unified perspective. This system consists of a server, a terminal, and a user.
[1135] First, the server collects data from public sources on the internet, such as government databases, news sites, and social networking services. This process utilizes APIs and web scraping techniques, using tools such as Python's requests library and BeautifulSoup. For example, it obtains policy information about COVID-19 from government open data portals and collects related articles from news sites.
[1136] The server then preprocesses the collected data, which includes removing special characters and whitespace, cleaning, and formatting, using the Python pandas library to organize the data into a unified format.
[1137] The preprocessed data is stored in a data management device by the server. The server uses a database system such as MySQL or PostgreSQL to efficiently store the data. The stored data is then input into a generative artificial intelligence model to train the model. Machine learning frameworks such as TensorFlow and PyTorch are used for this training.
[1138] A trained generative artificial intelligence model maps multiple viewpoints into a unified space. The server integrates data from different sources and viewpoints and calculates the spatial relationships between viewpoints using a clustering algorithm in Python's Scikit-learn library. As a result, the relationships between viewpoints are visualized, and important information is centralized.
[1139] A user requests specific information through a device, such as a PC or smartphone. An example of a prompt would be "Please tell me the latest measures against COVID-19." The device then sends this request to the server, which analyzes it using natural language processing (NLP) tools. The server uses Python's spaCy or NLTK to analyze the request, searches for relevant information in a unified space based on the analysis results, and generates comments or policy proposals using a generative artificial intelligence model if necessary.
[1140] Finally, the server sends the generated information to the device, which then displays the results to the user. For example, information such as "The latest government policy recommends vaccination in the ____ region. Furthermore, the self-isolation period has been changed from ____ to ____ days" can be provided. In this way, users can quickly obtain the latest and most reliable information.
[1141] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1142] Step 1: Gather public information
[1143] The server collects data from public sources on the Internet, including government databases, news sites, social networking services, etc. The server uses Python's requests library to retrieve the required data from APIs, and also uses BeautifulSoup to scrape articles from news sites.
[1144] Input: URL or API endpoint of a public source
[1145] Output: Raw data (text, JSON, etc.)
[1146] Step 2: Preprocessing the data
[1147] The server preprocesses the collected data, which includes removing unnecessary whitespace and special characters, standardizing the format, etc. The server uses the Python pandas library to clean the data and convert it into a unified format.
[1148] Input: Raw data
[1149] Output: Preprocessed data (unified format)
[1150] Step 3: Saving to the database
[1151] The server stores the preprocessed data in a data management device (database), such as MySQL or PostgreSQL. The server establishes a connection to the database and inserts data using SQL queries.
[1152] Input: Preprocessed data
[1153] Output: Data stored in the database
[1154] Step 4: Training the generative AI model
[1155] The server retrieves the latest data from the database and trains a generative AI model using machine learning frameworks such as TensorFlow and PyTorch. The server inputs the data into the model and trains it over multiple epochs.
[1156] Input: Data retrieved from the database
[1157] Output: A trained generative AI model
[1158] Step 5: Mapping to a unified space
[1159] The server uses a trained generative AI model to map multiple viewpoints into a unified space using a clustering algorithm using Python's Scikit-learn library, which calculates the spatial relationships between viewpoints obtained from different sources.
[1160] Input: A trained generative AI model and associated data
[1161] Output: viewpoint data mapped to a unified space
[1162] Step 6: Calculate the viewpoint position
[1163] The server calculates the relative positions of the viewpoints mapped to the unified space and prepares for visualization. It calculates the distance and relationship between viewpoints and generates a network graph. This process uses the Python networkx library.
[1164] Input: Mapped viewpoint data
[1165] Output: Data for visualizing the relative positions of viewpoints
[1166] Step 7: Receiving a user request
[1167] The user requests specific information through the device, for example by entering a prompt such as "Please tell me the latest information on COVID-19 countermeasures." The device then sends this request to the server.
[1168] Input: User request (prompt sentence)
[1169] Output: Request sent to the server
[1170] Step 8: Generate information
[1171] The server analyzes requests received from users using natural language processing (NLP) tools, such as Python's spaCy and NLTK. Based on the analysis results, the server searches for relevant information in a unified space and generates appropriate comments and policy proposals using a generative AI model.
[1172] Input: The request sent to the server
[1173] Output: Generated information (comments and policy suggestions)
[1174] Step 9: Provide information
[1175] The server sends the generated information to the device, which then displays the results to the user, such as, "The latest government policy recommends vaccination in the ____ region. In addition, the home quarantine period has been changed from ____ to ____ days."
[1176] Input: Generated information
[1177] Output: Information displayed on the user's terminal
[1178] In this way, specific actions are performed at each step, allowing the user to quickly obtain reliable, up-to-date information.
[1179] (Application example 1)
[1180] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1181] In today's world, businesses and individuals are constantly exposed to the threat of cyberattacks, resulting in increased risk of serious economic losses and information leaks. Cyberthreats evolve daily, and information about them is provided fragmentedly from numerous sources, making it difficult to unify, understand, and respond immediately. Furthermore, the lack of a system that can assess risks in real time and issue immediate alerts for high-risk threats means that prompt responses are delayed, potentially resulting in greater damage. To address these challenges and improve the effectiveness of cybersecurity, a system is needed that can unify the latest information, quickly and accurately assess risks, and notify users in real time.
[1182] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1183] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a database, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple perspectives into a unified space from the trained generative AI model, means for receiving requests from users, means for searching and generating information based on the requests, means for providing the generated information to users, means for risk assessing cybersecurity-related information, and means for issuing alerts to users about high-risk threats in real time. This enables rapid and accurate risk assessment of cyberattacks and the issuance of alerts in real time, thereby strengthening the security measures of companies and individuals.
[1184] "Public sources" refer to data sources that are accessible to anyone and from which information can be obtained, such as government databases, news sites, and social media.
[1185] "Data preprocessing" is the process of cleaning collected data, standardizing the format, removing unnecessary special characters, etc., to improve the quality of the data.
[1186] "Database" refers to a collection of data that stores collected and preprocessed data and structures it for easy access and retrieval.
[1187] A "generative AI model" is an artificial intelligence model that learns from collected data and has the ability to make inferences and predictions about new data.
[1188] A "unified space" refers to a data space that centrally aggregates data obtained from different sources and perspectives and presents it in a unified format and perspective.
[1189] A "clustering algorithm" is an algorithm used to group data based on similarity and calculate the relative positions of viewpoints.
[1190] "Request" means a request or inquiry sent by a User to a Server for specific information.
[1191] "Risk assessment" is the process of using collected data and trained generative AI models to assess the risk of potential cybersecurity-related threats and attacks.
[1192] "Real-time alerts" refer to warning messages that immediately notify users when cybersecurity threats increase and encourage them to take prompt action.
[1193] This invention is a system that trains a generative AI model using data collected from public sources and provides information from a unified perspective. The system functions through the following stages: [collection of public information], [data preprocessing], [storage in a database], [training of a generative AI model], [mapping to a unified space], [issuance of risk assessment and real-time alerts], and [provision of a user interface].
[1194] Hardware and Software Configuration
[1195] Hardware:
[1196] Server: Responsible for data collection, storage, model training, inference, and request processing.
[1197] Devices: Smartphones, smart glasses, head-mounted displays, robots, etc. Serve as the interface with the user.
[1198] software:
[1199] Data collection: API, web scraping tools (BeautifulSoup, Selenium)
[1200] Data preprocessing: Python (Pandas, NumPy)
[1201] Model training: TensorFlow, PyTorch
[1202] Clustering: Scikit-learn
[1203] Natural Language Processing: SpaCy, NLTK
[1204] User interface: Flutter (smartphones), Unity (smart glasses, head-mounted displays, robots)
[1205] Data collection and preprocessing
[1206] The server uses APIs and web scraping tools to collect data from public sources on the Internet (government databases, news sites, social media, etc.). The collected data is pre-processed using Python libraries to clean and standardize the format. The pre-processed data is then structured and stored in a database.
[1207] Generative AI models and mapping to a unified space
[1208] The saved data is input into a generative AI model, which is then trained. The generative AI model is developed using TensorFlow and PyTorch. The trained generative AI model maps the data obtained from multiple viewpoints into a unified space. The relative positions of the viewpoints are calculated and visualized using Scikit-learn's clustering algorithm.
[1209] Risk Assessment and Real-Time Alerts
[1210] The server evaluates cybersecurity-related risks based on the generated information and sends real-time alerts to users' systems about high-risk threats, allowing them to take prompt action.
[1211] User Interface
[1212] User requests are sent to the server via the device. The request is analyzed and relevant information is searched for and generated. User interfaces developed using Flutter and Unity are provided to users via smartphones, smart glasses, head-mounted displays, and robots.
[1213] Examples of concrete examples and prompts
[1214] For example, if a new ransomware attack is reported, the server collects relevant information, which is then scrutinized by the generative AI model and sent to the device, where it is notified to the user visually and audibly, providing immediate countermeasures.
[1215] Prompt Sentence Examples
[1216] User: "What's the latest ransomware attack?"
[1217] Security Assistant AI: "The latest ransomware attack is called ____ and occurred on ____ / ____ / __. Details of the attack are as follows: ____. As a countermeasure, we recommend ____."
[1218] This allows users to respond quickly to the latest cyber threats and strengthen their security.
[1219] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1220] Step 1:
[1221] The server uses APIs and web scraping tools (BeautifulSoup, Selenium, etc.) to collect data from public sources (government databases, news sites, social media, etc.). The input of this step is the public source, and the output is the collected raw data. Specifically, the server accesses the specified URL and retrieves the required information.
[1222] Step 2:
[1223] The server performs preprocessing on the collected raw data. Preprocessing includes data cleaning (removing unnecessary data and special characters), formatting unification (e.g., changing the date and time format), and text standardization (e.g., lowercasing). The input of this step is the collected raw data, and the output is preprocessed clean data. Specifically, the data is converted using Python's Pandas and NumPy libraries.
[1224] Step 3:
[1225] The server saves the preprocessed data in a database. The input of this step is clean data, and the output is data stored in the database. Specifically, it inserts and updates data in the database using SQL queries.
[1226] Step 4:
[1227] The server inputs the clean data into the generative AI model and trains the model. The input for this step is clean data retrieved from the database, and the output is a trained generative AI model. Specifically, the generative AI model is trained using TensorFlow or PyTorch.
[1228] Step 5:
[1229] The server uses a trained generative AI model to map multiple viewpoints into a unified space and calculates the relative positions of the viewpoints using Scikit-learn's clustering algorithm. The input to this step is the generative AI model and clean data, and the output is viewpoint data mapped into a unified space. Specifically, the clustering algorithm is applied to group the data.
[1230] Step 6:
[1231] The user sends an information request to the server through their device. The input in this step is the request from the user, and the output is the analysis result of the request content. Specifically, the user inputs the request using a smartphone or smart glasses.
[1232] Step 7:
[1233] Based on the received request, the server analyzes the request content using natural language processing (SpaCy or NLTK) and searches for related information from a unified space. The input to this step is the analysis result, and the output is the searched related information. Specifically, it extracts related information from a database based on the request content.
[1234] Step 8:
[1235] The server generates information using a generative AI model based on the searched related information and sends it to the terminal. The input of this step is the searched related information, and the output is the generated information. Specifically, the server generates information based on the generative AI model and sends it to the terminal.
[1236] Step 9:
[1237] The terminal provides the generated information received from the server to the user. The input of this step is the generated information from the server, and the output is information displayed and notified in a format that the user can understand. Specific operations include providing information to the user visually or audibly using a smartphone, smart glasses, a head-mounted display, or a robot.
[1238] Through this series of processing steps, the system alerts users in real time to high-risk cyber threats, enabling them to respond quickly.
[1239] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1240] This invention is a system that trains a generative AI model using data collected from public sources and provides information from a unified perspective. This system also incorporates an emotion engine that recognizes user emotions, improving the effectiveness of information provision. The system consists of a server, a terminal, and a user.
[1241] Public information collection and preprocessing
[1242] First, the server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, it obtains the necessary data using APIs and web scraping technology. For example, it collects the latest policy information on COVID-19 from the government's open data portal and related articles from news sites. Next, the server preprocesses the collected data. It cleans the data, standardizes the format, removes special characters and spaces, and arranges the data into a unified format.
[1243] Storing in a database and training the generated AI model
[1244] The preprocessed data is stored in a database by the server. The server then retrieves the latest data from the database and inputs it into the generative AI model. This generative AI model has the ability to learn from the collected data and make inferences and predictions about new data. The server trains the generative AI model so that it always reflects the latest information.
[1245] Mapping to a unified space and calculating the relative viewpoints
[1246] The server runs a technology that maps multiple viewpoints into a unified space using a trained generative AI model. During this process, data from different sources and viewpoints is centrally aggregated. Furthermore, a clustering algorithm is used to calculate and visualize the relative positions of each viewpoint. As a result, users can easily grasp the relevance and reliability of the information.
[1247] Emotion engine recognizes user emotions
[1248] The device sends user input and behavioral data (e.g., search history, clicks, input text, etc.) to the emotion engine, which analyzes the data and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.). This analysis is performed using natural language processing (NLP) and machine learning algorithms.
[1249] Receiving user requests and generating information
[1250] A user requests specific information through a device (such as a PC or smartphone). For example, they send a request such as, "Please tell me the latest measures against COVID-19." The device then sends this request to the server. The server receives the request from the user and analyzes it using natural language processing (NLP). Based on the analyzed request, the server searches for relevant information from a unified space and generates appropriate comments or policy proposals using a generative AI model.
[1251] Customize information based on user sentiment
[1252] Based on the analysis results from the emotion engine, the server adjusts the tone and content of the information generated by the generative AI model. For example, if the user is feeling anxious, the information provided will be generated in a reassuring tone. In this way, information is customized according to the user's emotional state, making it possible to provide more effective information.
[1253] Providing information and concrete examples
[1254] The server sends the generated information to the device, which then displays the results to the user. For example, if a user requests, "I want to know about COVID-19 countermeasures," the server generates the latest policy information and suggestions, and based on the results of the emotion engine, sends information such as, "The latest government policy recommends vaccination in the XX region. In addition, the self-isolation period has been changed from XX to XX days. If you are concerned, please refer to the detailed guidelines." to the device, which then displays this to the user. This system allows users to easily access the latest and most reliable information, which can be used to understand policies and support their daily lives. Furthermore, providing information based on the user's emotions increases user satisfaction and trust.
[1255] The processing flow will be explained below.
[1256] Step 1:
[1257] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.) by using REST APIs to retrieve data in JSON format or by using web scraping technology to obtain the required information.
[1258] Step 2:
[1259] The server pre-processes the data it receives, cleaning it (for example, removing whitespace and special characters), standardizing date and number formats, and arranging the data into a consistent format.
[1260] Step 3:
[1261] The server saves the pre-processed data to a database, where the information is stored efficiently and securely using INSERT operations into an SQL database.
[1262] Step 4:
[1263] The server retrieves the latest data from the database and inputs it into the generative AI model, which then uses the newly collected data to train the generative AI model.
[1264] Step 5:
[1265] The server trains the generative AI model, running machine learning algorithms that enable the model to make accurate predictions and inferences on new data.
[1266] Step 6:
[1267] The server maps multiple perspectives from the generative AI model into a unified space. Specifically, it plots data from different sources and perspectives in a multidimensional space and calculates their spatial relationships using a clustering algorithm.
[1268] Step 7:
[1269] The terminal sends the user's input data and behavioral data (e.g., past search history, click status, entered text, etc.) to the emotion engine.
[1270] Step 8:
[1271] The emotion engine analyzes data sent from the device and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.) using natural language processing (NLP) and machine learning algorithms.
[1272] Step 9:
[1273] Users request specific information through their device, for example, by typing in a request such as "Please tell me the latest information on COVID-19 countermeasures" and submitting it.
[1274] Step 10:
[1275] The terminal sends a request from the user to the server, and HTTP requests are generally used as the communication protocol.
[1276] Step 11:
[1277] The server analyzes the request received from the terminal, and uses natural language processing (NLP) to analyze the content of the request and identify its requirements.
[1278] Step 12:
[1279] The server searches for relevant information from the unified space based on the analysis results, leveraging the capabilities of SQL queries and generative AI models to extract the necessary information.
[1280] Step 13:
[1281] The server uses generative AI models to generate comments and policy proposals based on the retrieved information, using natural language generation (NLG) to provide the generated text in a format that is easy for users to understand.
[1282] Step 14:
[1283] The server adjusts the tone and content of the generated information based on the user's emotion analysis results obtained from the emotion engine. For example, if the user is feeling anxious, the information provided will be generated in a reassuring tone.
[1284] Step 15:
[1285] The server then sends the generated information to the terminal. Again, HTTP responses are generally used as the communication protocol.
[1286] Step 16:
[1287] The device displays the information received from the server to the user, either by updating a web page or displaying the information in the mobile app UI, making it easily accessible to the user.
[1288] This series of steps allows users to efficiently utilize public information and obtain the information they need quickly and accurately, as well as receive information tailored to their own emotions.
[1289] Example 2
[1290] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1291] In conventional information provision systems, it has been difficult to collect data from multiple information sources and provide users with reliable information from a unified perspective. Furthermore, the information provided does not take into account the user's emotional state, resulting in a decrease in user satisfaction. The present invention aims to solve these problems by handling data collected from a variety of information sources in a unified manner and providing information that corresponds to the user's emotional state.
[1292] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1293] In this invention, the server includes means for collecting data from public information sources, means for preprocessing the collected data, means for storing the preprocessed data in a database, means for inputting the stored data into a generative AI model, means for training the generative AI model, means for mapping multiple perspectives into a unified space from the trained generative AI model, means for analyzing emotions using user input data and behavioral data, means for receiving requests from users, means for searching and generating information based on the requests, means for providing the generated information to the user, and means for adjusting the tone and content of the generated information based on the results of the emotion analysis. This makes it possible to handle data collected from multiple information sources in a unified manner and provide appropriate information according to the user's emotional state.
[1294] "Public information sources" refer to information sources such as databases, news sites, and social media that are publicly available on the Internet.
[1295] "Means of collecting data" refers to the processes and techniques used to obtain the required information from public sources using APIs and web scraping techniques.
[1296] "Preprocessing" refers to the process of removing noise from collected data, standardizing the format, removing special characters and unnecessary spaces, and otherwise preparing the data for easier analysis.
[1297] "Means of database storage" refers to the process or technique by which pre-processed data is stored in a data management system.
[1298] A "generative AI model" refers to an artificial intelligence model that has the ability to learn from large amounts of data and make inferences and predictions based on new data.
[1299] "Training means" refers to the process or technique of inputting data into a generative AI model and optimizing the model's parameters.
[1300] "Means of mapping to a unified space" refers to the process of centrally aggregating data from multiple perspectives obtained by a generative AI model and visualizing the relationships between the perspectives in a low-dimensional space.
[1301] "Means of emotion analysis" refers to natural language processing and machine learning technologies that recognize users' emotional states based on their input data and behavioral data.
[1302] "Means for receiving requests from users" refers to the process by which the server receives inquiries or information requests sent by users through their terminals.
[1303] "Means for information retrieval and generation" refers to the process of analyzing a user's request, searching for relevant information from a database based on that request, and providing appropriate information using a generative AI model.
[1304] "Means for providing information to a user" refers to the process of displaying the generated information to a user through a terminal.
[1305] "Measures to adjust tone and content" refers to the process of changing the expression and content of generated information based on the results of sentiment analysis in accordance with the user's emotional state.
[1306] This system trains a generative AI model using data collected from public sources to provide information from a unified perspective. It also incorporates an emotion engine that recognizes the user's emotional state and provides information accordingly. The system consists of a server, a terminal, and a user.
[1307] Public information collection and preprocessing
[1308] The server collects data from public sources on the Internet (government databases, news sites, social media, etc.). Specifically, the server obtains data using APIs and web scraping technologies (e.g., BeautifulSoup or Scrapy). For example, the server collects the latest policy information about COVID-19 from the government's open data portal and related articles from news sites. The server then preprocesses the collected data. This involves removing noise, standardizing the format, and removing special characters and spaces, thereby arranging the data into a unified format. This makes it easier to input into the generative AI model.
[1309] Storing in a database and training the generated AI model
[1310] The preprocessed data is stored in a database by the server. Specifically, the server connects to a database system such as MySQL or MongoDB and stores the data. The server then retrieves the latest data from the database and inputs it into a generative AI model (e.g., GPT-3 or BERT). This generative AI model has the ability to learn from the collected data and make inferences and predictions based on new data. The server trains the generative AI model using libraries such as TensorFlow and PyTorch to ensure that it always reflects the latest information.
[1311] Mapping information to a unified space and calculating spatial relationships
[1312] The server maps multiple perspectives obtained from the trained generative AI model into a unified space. In this process, data obtained from different information sources and perspectives is centrally aggregated. The server then uses a clustering algorithm (e.g., k-means) to calculate the relative positions of each perspective and visualize them in a low-dimensional space. This allows users to easily grasp the relevance and reliability of the information.
[1313] Emotion engine recognizes user emotions
[1314] The device sends the user's input data and behavioral data (e.g., search history, clicks, input text, etc.) to the emotion engine. The emotion engine analyzes this data and recognizes the user's current emotional state (joy, sadness, anger, fear, etc.). Specifically, it performs the analysis using natural language processing (NLP) and machine learning algorithms (e.g., the emotion analysis model VADER).
[1315] Receiving user requests and generating information
[1316] A user requests specific information through a device (e.g., a PC or smartphone). For example, the user types, "Please tell me the latest measures against COVID-19." This request is sent from the device to the server. The server receives the request and analyzes it using natural language processing (NLP) technology. Based on the analysis results, the server searches for relevant information from a unified space and generates appropriate information using a generative AI model.
[1317] Regulating and providing information based on emotions
[1318] Based on the analysis results from the emotion engine, the server adjusts the tone and content of the generated information. For example, if the user is feeling anxious, the server will provide information in a reassuring tone, such as, "Don't worry, the latest government policy recommends vaccination in the XX area." The final information is sent from the server to the device, which then displays it to the user. This allows users to always have easy access to the latest and most reliable information, allowing them to use the information with greater peace of mind.
[1319] Examples of concrete examples and prompts
[1320] For example, if you receive a request such as "I want to know the latest information on COVID-19 countermeasures," the prompt might look like this:
[1321] What is the latest information on COVID-19 prevention measures? My current emotional state is anxiety.
[1322] By inputting this prompt into the generative AI model, appropriate information is provided.
[1323] This system provides more effective information provision by customizing information according to the user's emotional state, and by centralizing data collected from various sources, users can quickly obtain reliable information.
[1324] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1325] Step 1:
[1326] The server collects data from public sources (government databases, news sites, social media, etc.). Specifically, the server obtains information using APIs or web scraping technologies (e.g., BeautifulSoup or Scrapy). The input is the URL or API key of the public source, and the output is the collected raw data. For example, the server obtains the latest policy information about COVID-19 from the government's open data portal.
[1327] Step 2:
[1328] The server preprocesses the collected data. The input is the raw data collected in step 1, and the output is clean, uniformly formatted data. Specifically, the server uses regular expressions to remove noise, standardize the format, and remove special characters and whitespace. This prepares the data in a form that is easy to input into the generative AI model.
[1329] Step 3:
[1330] The server saves the preprocessed data in a database. The input is the preprocessed data, and the output is the data stored in the database. Specifically, the server connects to a database system such as MySQL or MongoDB and executes SQL statements to insert data. This allows the data to be managed in an organized manner.
[1331] Step 4:
[1332] The server retrieves the latest data from the database and inputs it into the generative AI model. The input is the formatted data retrieved from the database, and the output is the data processed by the generative AI model. Specifically, the server retrieves the data and inputs it into the generative AI model (e.g., GPT-3 or BERT) using libraries such as TensorFlow or PyTorch.
[1333] Step 5:
[1334] The server trains the generative AI model. The input is the training data and the model's initial parameters, and the output is the trained model. For example, the server uses a dataset to optimize the model's parameters and improve its inference ability on new data. This training process is performed using libraries such as TensorFlow and PyTorch.
[1335] Step 6:
[1336] The server maps multiple viewpoints from a trained generative AI model into a unified space. The input is viewpoint data generated by the model, and the output is a unified viewpoint space. Specifically, the server visualizes the viewpoint data in a low-dimensional space using UMAP (Uniform Manifold Approximation and Projection) or a clustering algorithm (e.g., k-means).
[1337] Step 7:
[1338] The device sends the user's input data and behavioral data to the emotion engine. The input is the user's input data (search history, clicks, input text, etc.), and the output is the emotion analysis results. The device sends this data to the emotion engine, which then analyzes it using natural language processing technology.
[1339] Step 8:
[1340] A user requests specific information through a device. The input is the user's request (e.g., "Please tell me the latest measures against COVID-19"), and the output is the request data. This request is sent from the device to the server.
[1341] Step 9:
[1342] The server receives the user's request and analyzes it using natural language processing technology. The input is the text data of the user's request, and the output is the analysis result. Based on the analysis result, the server searches for relevant information from the unified space and generates appropriate information using a generative AI model.
[1343] Step 10:
[1344] The server adjusts the tone and content of the generated information based on the analysis results from the emotion engine. The input is the emotion analysis result and the generated information, and the output is the adjusted information. For example, if the user is feeling anxious, the server will provide information in a reassuring tone.
[1345] Step 11:
[1346] The server sends the generated information to the terminal, which then displays it to the user. The input is the adjusted information, and the output is the information displayed to the user. The terminal displays on the user's smartphone or PC, "The latest government policy recommends vaccination in the XX area. Additionally, the home quarantine period has been changed from XX to XX days."
[1347] (Application example 2)
[1348] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1349] In modern society, the importance of security information is increasing day by day. However, there is still a lack of systems that allow users to obtain reliable security information in real time and provide that information tailored to the user's emotional state. Furthermore, there is the challenge of properly recognizing users' emotions, such as anxiety and fear, and providing customized information based on those emotions. Therefore, it is necessary to provide a system that allows users to receive security information with peace of mind.
[1350] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1351] In this invention, the server includes means for collecting data from public information sources, means for pre-processing the collected data, and means for storing the pre-processed data in a database, thereby enabling security information collected from public information sources to be provided to the user in real time and further customized by recognizing the user's emotional state.
[1352] "Public sources" are sources that are free or publicly available on the internet and can be accessed by anyone. Examples include government databases, news sites, and social media.
[1353] "Methods of data collection" refers to technologies and tools used to automatically obtain the required data from public sources on the Internet, including the use of APIs and web scraping techniques.
[1354] "Preprocessing methods" refers to techniques for improving the quality of collected data and arranging it into a unified format, such as data cleaning, formatting standardization, and removal of special characters.
[1355] "Means of storing data in a database" refers to a system that efficiently stores pre-processed data and manages it for easy retrieval and use. Examples include RDBMS and NoSQL databases.
[1356] "Means of input to generative AI models" refers to technologies that provide data stored in a database to generative AI models for use in training and prediction.
[1357] "Training methods" refers to the techniques used to train generative AI models with the latest data and improve their performance, including machine learning algorithms and data feeds.
[1358] "Means of mapping to a unified space" refers to techniques that unify data from different sources and perspectives to allow users to understand the relevance and reliability of the information. Specifically, this includes dimensionality reduction and clustering algorithms.
[1359] "Means for receiving requests" refers to the technology used to obtain information requests from users and communicate them to the system, including the user interface and the NLP engine.
[1360] "Information search and generation means" refers to technologies that search for relevant information from databases based on user requests and generate appropriate answers or comments, including search algorithms and generative AI models.
[1361] "Means for providing to users" refers to technologies for presenting generated information to users in an easy-to-understand manner, including user interfaces and notification systems.
[1362] "Emotional state recognition" refers to technologies for identifying a user's current emotional state based on their input and behavioral data, including natural language processing and machine learning algorithms.
[1363] "Information tailoring measures" refers to techniques for adjusting the tone and content of information provided in response to the perceived emotional state of the user, including tone modification and information filtering.
[1364] This invention is a system that collects data from public sources, uses a generative AI model to generate information requested by users, and further customizes the information according to the user's emotional state. The system mainly consists of three elements: a server, a terminal, and a user.
[1365] System configuration
[1366] server
[1367] The server has the following main functions:
[1368] Data collection from public sources: The server collects data in real time from public sources on the Internet, using APIs and web scraping techniques to extract necessary security information from government databases, news sites, social media, etc.
[1369] Data preprocessing: The collected data is preprocessed by data cleaning, standardizing the format, removing special characters, etc.
[1370] Database storage: The preprocessed data is stored in a database that can be efficiently searched and updated, typically using an RDBMS or NoSQL database.
[1371] Information generation using a generative AI model: The stored data is input into a generative AI model to train the model. The trained generative AI model is then used to generate information based on user requests.
[1372] Emotion recognition: Analyzes user input data (e.g., search history, clicks, input text) and uses an emotion engine to recognize the user's current emotional state.
[1373] Information customization: Adjusting the tone and content of generated information based on perceived emotional state.
[1374] Terminal
[1375] The device (e.g., smartphone or PC) has the following features:
[1376] Sending a user request: A user requests specific information and sends it to the server via their device. For example, a request could be, "Please tell me about the latest cyber-attack information."
[1377] Receiving and displaying information: Receives information generated by the server and displays it to the user. The display method can be in various formats such as text or notification.
[1378] user
[1379] The user does the following:
[1380] Information Request: Using a smartphone or PC, a user requests specific information. This request is made in natural language and sent to the server.
[1381] Review the information: Review the information received and take further action as necessary.
[1382] Specific examples of processing
[1383] When a user requests, "I'd like to know about recent cyber-attack trends," the server first collects the latest security-related data from public sources. This data is preprocessed and stored in a database. Then, a generative AI model is used to generate appropriate information and customize it based on the user's emotional state. Finally, the generated, specific answer (e.g., "Recent cyber-attacks have been characterized by large-scale phishing campaigns and the emergence of new ransomware. In particular, there has been an increase in techniques that identify and encrypt sensitive data within companies and demand ransom. Effective defensive measures include regular software updates and strengthened employee education. If you have concerns, please refer to our more detailed security policy") is provided to the user via their device.
[1384] Prompt Sentence Examples
[1385] Examples of specific prompts based on user input include:
[1386] User input: "I'd like to know about recent trends in cyber attacks."
[1387] Prompt: "User's feelings: anxious. Collect the security-related information: (Recently collected security information). User's request: I would like to know about recent cyber-attack trends."
[1388] This specific example allows users to obtain the latest and most reliable security information in real time, while also providing specific measures to alleviate their concerns.
[1389] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1390] Step 1:
[1391] The server collects data from public sources. Specifically, it uses APIs and web scraping technology to obtain real-time data from government databases, news sites, and social media on the Internet. The collected data is temporarily stored as raw data. Inputs include the URL and API key of the target public source, and the obtained raw data is included as output.
[1392] Step 2:
[1393] The server preprocesses the collected data. Specific operations include data cleaning (removing missing values and outliers), standardizing formats (e.g., standardizing character codes and date formats), removing special characters, etc. The input data is the raw data collected in the previous step, and the output data is the preprocessed clean data.
[1394] Step 3:
[1395] The server stores the preprocessed data in a database. Specifically, it uses an RDBMS or NoSQL database to efficiently store data. The input data is the preprocessed clean data, and the output is the storage data stored in the database.
[1396] Step 4:
[1397] The server inputs the stored data into the generative AI model and trains the model. It retrieves the pre-processed and stored data from the database and supplies it to the generative AI model, updating the model with the latest information and training it. The input data is the data retrieved from the database, and the output data is the trained generative AI model.
[1398] Step 5:
[1399] The server maps multiple perspectives from the trained generative AI model into a unified space. Specifically, it aggregates data from different information sources and perspectives into a unified space using a clustering algorithm and visualizes it. The input data is the trained generative AI model, and the output data is the mapped perspective data.
[1400] Step 6:
[1401] The server receives requests from users and analyzes them. When a user sends a request from a device, the server analyzes the request using natural language processing (NLP). The input data is the user's request, and the output data is the analysis result.
[1402] Step 7:
[1403] The server searches and generates information based on the analyzed request, searches for relevant information from the database, and generates an appropriate answer using a generative AI model. The input data is the analysis result and the storage data saved in the database, and the output data is the generated answer.
[1404] Step 8:
[1405] The server recognizes the user's emotional state. It uses an emotion engine to analyze the user's emotional state (e.g., relief, anxiety, interest, etc.) from their search history and input text. The input data is the user's behavioral data, and the output data is the recognized emotional state.
[1406] Step 9:
[1407] The server adjusts the generated information based on the recognized emotional state, specifically by changing the tone of the text or filtering the information. The input data is the recognized emotional state and the generated response, and the output data is the adjusted response.
[1408] Step 10:
[1409] The terminal provides the generated information received from the server to the user. The terminal presents the information to the user in text format, notification format, etc. The input data is the adjusted response from the server, and the output data is the information received by the user.
[1410] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1411] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1412] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1413] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1414] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1415] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1416] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1417] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1418] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1419] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1420] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1421] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1422] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1423] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1424] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1425] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1426] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1427] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1428] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1429] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1430] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1431] The following is further disclosed regarding the above embodiment.
[1432] (Claim 1)
[1433] means of collecting data from public sources;
[1434] means for pre-processing the collected data;
[1435] means for storing the preprocessed data in a database;
[1436] a means for inputting the stored data into a generative AI model;
[1437] A means for training a generative AI model; and
[1438] A means of mapping multiple viewpoints into a unified space from a trained generative AI model; and
[1439] a means for receiving requests from users;
[1440] a means for retrieving and generating information based on the request;
[1441] a means for providing the generated information to a user;
[1442] A system including:
[1443] (Claim 2)
[1444] The system of claim 1, wherein data is collected from government databases, news sites, and social media as public information sources.
[1445] (Claim 3)
[1446] The system of claim 1, wherein a clustering algorithm is used to calculate the positional relationship of each viewpoint in mapping to a unified space.
[1447] (Claim 4)
[1448] 10. The system of claim 1, wherein the user request is analyzed using natural language processing (NLP).
[1449] (Claim 5)
[1450] 10. The system of claim 1, wherein the generated information is generated using natural language generation (NLG).
[1451] "Example 1"
[1452] (Claim 1)
[1453] means of collecting data from public sources;
[1454] means for pre-processing the collected data;
[1455] means for storing the preprocessed data in a data management device;
[1456] means for inputting the stored data into a generative artificial intelligence model;
[1457] means for training a generative artificial intelligence model;
[1458] a means for mapping multiple viewpoints into a unified space from a trained generative artificial intelligence model;
[1459] means for calculating positional relationships between viewpoints obtained from different sources;
[1460] a means for receiving requests from users;
[1461] a means for retrieving and generating information based on the request;
[1462] a means for providing the generated information to a user;
[1463] A system including:
[1464] (Claim 2)
[1465] 10. The system of claim 1, wherein the data is collected from public sources such as government databases, news sites, and social networking services.
[1466] (Claim 3)
[1467] The system of claim 1, wherein a grouping algorithm is used to calculate the positional relationship of each viewpoint in mapping to a unified space.
[1468] "Application Example 1"
[1469] (Claim 1)
[1470] means of collecting data from public sources;
[1471] means for pre-processing the collected data;
[1472] means for storing the preprocessed data in a database;
[1473] a means for inputting the stored data into a generative AI model;
[1474] A means for training a generative AI model; and
[1475] A means of mapping multiple viewpoints into a unified space from a trained generative AI model; and
[1476] a means for receiving requests from users;
[1477] a means for retrieving and generating information based on the request;
[1478] a means for providing the generated information to a user;
[1479] A means of risk-assessing cybersecurity-related information;
[1480] A means to alert users to high-risk threats in real time; and
[1481] A system including:
[1482] (Claim 2)
[1483] The system of claim 1, wherein data is collected from government databases, news sites, and social media as public information sources.
[1484] (Claim 3)
[1485] The system of claim 1, wherein a clustering algorithm is used to calculate the positional relationship of each viewpoint in mapping to a unified space.
[1486] "Example 2: Combining Emotion Engines"
[1487] (Claim 1)
[1488] means of collecting data from public sources;
[1489] means for pre-processing the collected data;
[1490] means for storing the preprocessed data in a database;
[1491] a means for inputting the stored data into a generative AI model;
[1492] A means for training a generative AI model; and
[1493] A means of mapping multiple viewpoints into a unified space from a trained generative AI model; and
[1494] A means for analyzing emotions using user input data and behavioral data;
[1495] a means for receiving requests from users;
[1496] a means for retrieving and generating information based on the request;
[1497] a means for providing the generated information to a user;
[1498] a means for adjusting the tone and content of the generated information based on the results of the sentiment analysis;
[1499] A system including:
[1500] (Claim 2)
[1501] 10. The system of claim 1, wherein the system collects data from government databases, news sites, and social media.
[1502] (Claim 3)
[1503] The system of claim 1, wherein a clustering algorithm is used to calculate the positional relationship of each viewpoint in mapping to a unified space.
[1504] "Application example 2 when combining emotion engines"
[1505] (Claim 1)
[1506] means of collecting data from public sources;
[1507] means for pre-processing the collected data;
[1508] means for storing the preprocessed data in a database;
[1509] a means for inputting the stored data into a generative AI model;
[1510] A means for training a generative AI model; and
[1511] A means of mapping multiple viewpoints into a unified space from a trained generative AI model; and
[1512] a means for receiving requests from users;
[1513] a means for retrieving and generating information based on the request;
[1514] a means for providing the generated information to a user;
[1515] a means for recognizing the emotional state of a user;
[1516] means for adjusting the generated information based on the emotional state of the user;
[1517] A system including:
[1518] (Claim 2)
[1519] The system of claim 1, wherein data is collected from government databases, news sites, and social media as public information sources.
[1520] (Claim 3)
[1521] The system of claim 1, wherein a clustering algorithm is used to calculate the positional relationship of each viewpoint in mapping to a unified space.
[1522] (Claim 4)
[1523] 10. The system of claim 1, wherein the user's emotional state is analyzed using natural language processing and machine learning algorithms.
[1524] (Claim 5)
[1525] 10. The system of claim 1, wherein the system provides security information collected from public sources to a user in real time. [Explanation of symbols]
[1526] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means of collecting data from public sources; means for pre-processing the collected data; means for storing the preprocessed data in a database; a means for inputting the stored data into a generative AI model; A means for training a generative AI model; and A means of mapping multiple viewpoints into a unified space from a trained generative AI model; and a means for receiving requests from users; a means for retrieving and generating information based on the request; a means for providing the generated information to a user; A system including:
2. The system of claim 1 , wherein data is collected from government databases, news sites, and social media as public information sources.
3. The system according to claim 1, wherein a clustering algorithm is used to calculate the positional relationship of each viewpoint in mapping to the unified space.
4. 10. The system of claim 1, wherein the user request is analyzed using natural language processing (NLP).
5. 10. The system of claim 1, wherein the generated information is generated using natural language generation (NLG).
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A