system

The system addresses the inefficiencies in collecting and summarizing technical literature and intellectual property information by using a database and generative AI to automatically preprocess and generate summaries, enhancing efficiency and reducing time consumption.

JP2026063744APending Publication Date: 2026-04-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-01
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Small-scale enterprises and individuals face challenges in efficiently collecting and summarizing vast amounts of technical literature and intellectual property information, which is time-consuming and labor-intensive.

Method used

A system equipped with a database for organizing data, preprocessing capabilities, and a generative artificial intelligence model to collect, preprocess, and generate summaries of technical literature and intellectual property information based on user queries.

Benefits of technology

Enables users to efficiently gather and understand the latest technology trends and patent information, significantly reducing the effort and time required for information gathering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063744000001_ABST
    Figure 2026063744000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 Means for collecting technical documents and intellectual property information from a database, Means for preprocessing the collected data and removing irrelevant information from the text, Means for training a generative artificial intelligence model using the preprocessed data, Means for receiving a search query from a user, Means for extracting relevant information from a database based on the received search query, Means for generating a summary using the extracted information with a generative artificial intelligence model, Means for providing the generated summary to the user, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the modern business environment, it is extremely important to quickly and accurately obtain information related to the latest technologies and intellectual property. However, it is very time-consuming and labor-intensive to find and summarize relevant information from a vast amount of technical documents and patent information. In particular, small-scale enterprises and individuals currently lack means to efficiently collect and analyze such information, which is an issue.

Means for Solving the Problems

[0005] To solve the above problems, the present invention provides a system equipped with means for collecting technical literature and intellectual property information from a database and preprocessing them. Furthermore, it includes means for training a generative artificial intelligence model using the preprocessed data, and for generating and providing summaries based on search queries from users using this model. This system allows users to efficiently collect relevant information and quickly grasp the latest technology trends and patent information, thereby significantly reducing the effort and time required for information gathering.

[0006] A "database" is a system that structures and organizes a collection of data, and is a means of efficiently searching, storing, and managing related information.

[0007] "Technical literature" refers to documents containing technical information and research results, such as academic papers, scholarly books, and technical reports.

[0008] "Intellectual property information" refers to information related to intellectual property rights such as patents, trademarks, copyrights, and design rights.

[0009] "Collection" is the process of obtaining, integrating, or storing specific data.

[0010] "Preprocessing" refers to the process of converting collected data into a format that can be applied to analysis and model training, and includes noise reduction and format conversion.

[0011] A "generative artificial intelligence model" is a machine learning model that automatically generates new information from data for the purpose of prediction and generation.

[0012] "Training" is the process by which a model learns data patterns and improves itself.

[0013] A "user" is an individual or organization that uses the system to obtain services or information.

[0014] A "search query" is a keyword or phrase that a user inputs into a system to search for specific information.

[0015] "Extraction" is the process of retrieving necessary information from a database or dataset.

[0016] A "summary" is a document that briefly summarizes the main points of the original information.

[0017] "Provision" is the act of delivering the generated information or service to the user.

[0018] A "system" is a collection of devices or programs in which multiple elements operate in cooperation to provide specific functions or services.

Brief Description of the Drawings

[0019] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0020] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be described.

[0022] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0023] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0024] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0025] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0027] [First Embodiment]

[0028] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0029] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0032] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0035] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0039] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0040] This invention relates to a system for collecting technical literature and intellectual property information, enabling users to efficiently obtain this information. The system is characterized by collecting information from a database, performing preprocessing, generating a summary using a generative artificial intelligence model, and providing it to the user.

[0041] System programming and processing

[0042] 1. Data Collection

[0043] The server connects to technical literature databases and intellectual property databases to collect relevant information.

[0044] The server queries the database using keywords such as "autonomous driving" and retrieves the corresponding data.

[0045] The server stores the retrieved data in a structured format such as JSON.

[0046] 2. Data preprocessing

[0047] The server converts the collected data into text format.

[0048] The server cleans up the data by removing noise from the text (e.g., HTML tags and special characters).

[0049] The server tokenizes the clean data and removes irrelevant stop words.

[0050] The server saves the pre-processed data to the database.

[0051] 3. Training the AI ​​model

[0052] The server uses the pre-processed dataset to train a generative artificial intelligence model.

[0053] The server uses generative artificial intelligence models (e.g., BERT or GPT) to learn important patterns in the data.

[0054] The server stores the trained model and deploys it for inference.

[0055] 4. Acceptance of user queries

[0056] The terminal receives search queries from the user through the input interface.

[0057] The user enters keywords such as "latest information on autonomous driving technology."

[0058] 5. Query analysis and extraction of corresponding data

[0059] The server receives the search query sent from the terminal and parses the keywords.

[0060] The server extracts relevant technical literature and intellectual property information from the database based on the query.

[0061] 6. Summary Generation

[0062] The server inputs the extracted data into a generative artificial intelligence model and generates a summary.

[0063] The server verifies the summary results to ensure that the main points are reflected.

[0064] 7. Providing a summary

[0065] The server converts the generated summary into JSON format and sends it to the terminal.

[0066] The device displays summary information to the user.

[0067] Specific example

[0068] Data Acquisition and Preprocessing

[0069] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to autonomous driving. The collected data is stored in JSON format, and after noise reduction and tokenization, it is stored in the database.

[0070] Model training and query analysis

[0071] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize important information related to autonomous driving technology. When a user enters "latest information on autonomous driving technology" as a query, the server parses this query and extracts relevant papers and patents.

[0072] Summary generation and delivery

[0073] The extracted information is summarized by a generative artificial intelligence model. For example, a summary might be generated stating, "The latest autonomous driving technologies include AI-powered real-time road surface detection and high-precision obstacle detection using laser sensors." The generated summary is sent from the server to the terminal, where the user can view it through the interface.

[0074] Based on the above, the present invention is a system that enables users to effectively collect and summarize technical documents and intellectual property information.

[0075] The following describes the processing flow.

[0076] Step 1:

[0077] The server connects to technical literature databases and intellectual property databases. For example, it uses APIs to collect data from the "technical literature database" and the "intellectual property database."

[0078] Step 2:

[0079] The server converts the collected data into JSON format and saves it to local or cloud storage. The collected data includes the title of the paper, authors, abstracts, patent numbers, and summaries of the inventions.

[0080] Step 3:

[0081] The server preprocesses the collected data. Specifically, it removes HTML tags and special characters from the text. Next, it tokenizes the cleansed data and removes unwanted words (stop words).

[0082] Step 4:

[0083] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model to begin the learning process.

[0084] Step 5:

[0085] The server evaluates the trained generative artificial intelligence model on a test dataset to verify its accuracy. Once the evaluation is complete, the model is deployed for inference.

[0086] Step 6:

[0087] The terminal receives search queries from users through its interface. For example, it might accept a query such as "latest trends in autonomous driving technology."

[0088] Step 7:

[0089] The server analyzes the search query received from the terminal and extracts keywords. For example, "autonomous driving technology" and "latest trends" might be extracted as keywords.

[0090] Step 8:

[0091] The server extracts relevant technical literature and intellectual property information from the database based on the analyzed query. The extracted data is temporarily loaded into memory.

[0092] Step 9:

[0093] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The summary includes key technical points, findings, and application examples.

[0094] Step 10:

[0095] The server converts the generated summary into JSON format and sends it to the terminal. Because the transmitted summary data is structured, it is displayed in an easy-to-use format.

[0096] Step 11:

[0097] The terminal displays the received summary results in the user interface. Users can review the summary information corresponding to their search queries and quickly obtain the necessary technical information.

[0098] This series of processes allows users to efficiently collect the latest technical literature and intellectual property information, and obtain the necessary information in a summarized format.

[0099] (Example 1)

[0100] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0101] Because technical literature and intellectual property information are vast, it is difficult for individual users to efficiently collect the necessary information and summarize and understand the key points. Traditional systems primarily rely on manual information gathering and summarization, which is time-consuming and labor-intensive. Furthermore, the process of extracting and summarizing appropriate information based on search queries is often inefficient. Therefore, there is a need for a system that can automatically collect, preprocess, and summarize technical literature and intellectual property information, and efficiently provide it to users.

[0102] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0103] In this invention, the server includes means for collecting relevant information from a technical literature database and an intellectual property database; means for converting the collected information into text format, removing noise, and creating clean data; and means for tokenizing the clean data and preprocessing irrelevant information. This makes it possible to automatically collect and preprocess technical literature and intellectual property information to generate a dataset.

[0104] The server also includes means for training a generative artificial intelligence model using a preprocessed dataset, means for receiving search queries from users, means for analyzing the received search queries and extracting relevant information from technical literature databases and intellectual property databases, means for generating summaries of the extracted information using the generative artificial intelligence model, and means for converting the generated summaries into JSON format and providing them to users. This makes it possible to efficiently extract relevant information, summarize the key points, and provide them to users.

[0105] A "technical literature database" is a database that collects technical information, such as academic papers and technical reports, and stores it in a searchable format.

[0106] An "intellectual property database" is a database that collects and stores information related to intellectual property, such as patents, trademarks, and copyrights, in a searchable format.

[0107] A "generative artificial intelligence model" is an artificial intelligence model that has the ability to generate new data in tasks such as natural language processing and image generation.

[0108] "JSON format" is an abbreviation for JavaScript (registered trademark) Object Notation, and is a lightweight data exchange format for structuring and storing data.

[0109] "Noise reduction" is the process of removing unnecessary information (e.g., HTML tags and special characters) from data to generate clean data.

[0110] A "query" is a type of inquiry that a user enters in a database or search system to retrieve specific information.

[0111] "Tokenization" is the process of dividing text data into semantic units (tokens) such as words and phrases.

[0112] "Preprocessing" refers to the process of preparing data, such as cleaning up and standardizing its format, before analyzing the data or training models.

[0113] A "summary" is information that extracts the important points from a vast amount of information and presents them in a short format.

[0114] "Analysis" is the process of analyzing data and query content to understand and process their meaning and structure.

[0115] "Extraction" is the process of retrieving data or information based on specific conditions.

[0116] This invention relates to a system for efficiently collecting and processing technical literature and intellectual property information and providing it to users in a summarized format. The system collects relevant information from technical literature databases and intellectual property databases, performs preprocessing, generates summaries using a generative artificial intelligence model, and provides them to users.

[0117] Data collection

[0118] The server periodically accesses technical literature databases (e.g., IEEE Xplore) and intellectual property databases (e.g., USPTO) to collect relevant information. For example, the server automatically sends queries at 3:00 AM every day to retrieve the latest papers and patent information related to keywords such as "autonomous driving" and "AI technology." The collected information is temporarily stored in data storage in a structured JSON format.

[0119] Data preprocessing

[0120] The server converts the collected data into text format and removes noise. First, HTML tags and special characters are removed using regular expressions. Next, the text data is tokenized (divided into words and phrases), and irrelevant stop words (e.g., "a" and "the" in English) are removed using the NLTK library. The pre-processed data is then saved back to the database.

[0121] Training an AI model

[0122] The server trains a generative artificial intelligence (AI) model using a pre-processed dataset. For example, it trains a BERT model using the Hugging Face Transformers library. The server uses 80% of the data for training and 20% for validation. Once the training is complete, the model is stored in the Hugging Face model hub.

[0123] Acceptance of user queries

[0124] The terminal receives search queries from the user through an input interface. For example, it might be implemented as a web application running in a browser, where the user enters a query such as "latest autonomous driving technology" into the search bar. This query is sent to the server in real time.

[0125] Query analysis and extraction of corresponding data

[0126] The server receives search queries sent from the terminal and analyzes the keywords using TF-IDF (Term Frequency-Inverse Document Frequency). Based on the analysis results, the server extracts relevant information from the database. For example, it might pick out literature and patents related to "latest autonomous driving technology."

[0127] Summary generation

[0128] The server inputs the extracted literature and patent information into a generative artificial intelligence model to generate a summary. Specifically, the extracted text is fed into a BERT model to generate a summarized text. At this time, the model undergoes multiple validations to ensure that important points are included.

[0129] Summary

[0130] The server converts the generated summary into JSON format and sends it to the terminal. The terminal displays this summary information to the user. For example, in the user interface of a web application, the summary text is displayed in an easy-to-read format. The user can review the displayed summary and access detailed information as needed.

[0131] Specific example

[0132] Data Acquisition and Preprocessing

[0133] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to "autonomous driving." The collected data is stored in JSON format and, after preprocessing such as removing HTML tags and special characters and tokenizing the text, is stored in the database.

[0134] Model training and query analysis

[0135] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize key information related to "autonomous driving technology." When a user enters "latest information on autonomous driving technology" as a query, the server analyzes this query and extracts relevant papers and patents.

[0136] Summary generation and delivery

[0137] The extracted information is summarized by a generative artificial intelligence model. For example, a summary might be generated stating, "The latest autonomous driving technologies include AI-powered real-time road surface detection and high-precision obstacle detection using laser sensors." The generated summary is sent from the server to the terminal, where the user can view it in a web browser.

[0138] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0139] Step 1: Data collection from the database

[0140] The server connects to technical literature databases and intellectual property databases to collect information related to specific keywords (e.g., "autonomous driving" or "AI technology"). The input is a keyword query, and the output is the collected data in JSON format. Specifically, the server automatically sends queries at 3 AM every day and temporarily stores the retrieved papers and patent information in data storage.

[0141] Step 2: Data Preprocessing

[0142] The server converts the collected data into text format and removes noise. The input is collected data in JSON format, and the output is clean text data. Specifically, the server uses regular expressions to remove HTML tags and special characters, then tokenizes the text data into words and phrases, and removes irrelevant stop words using the NLTK library. This clean data is then stored again in the database.

[0143] Step 3: Training the AI ​​model

[0144] The server trains a generative artificial intelligence model using a preprocessed dataset. The input is a preprocessed dataset, and the output is a trained generative AI model. Specifically, the server trains a BERT model using the Hugging Face Transformers library, using 80% of the data for training and 20% for validation. The trained model is stored in the Hugging Face model hub.

[0145] Step 4: Accepting User Queries

[0146] The terminal receives search queries from the user through an input interface. The input is the user's search query (e.g., "latest autonomous driving technology"), and the output is that query being sent to the server. Specifically, the user enters the query into the search bar of the web application, and that query is sent to the server in real time.

[0147] Step 5: Query analysis and extraction of corresponding data

[0148] The server receives search queries sent from terminals and parses the keywords. The input is the user's search query, and the output is relevant information extracted from the database. Specifically, the server uses TF-IDF (Term Frequency-Inverse Document Frequency) to parse the query and extracts relevant technical documents and patent information from the database based on the results.

[0149] Step 6: Summary Generation

[0150] The server inputs the extracted information into a generative artificial intelligence model to generate a summary. The input includes extracted technical documents and patent information, and the output is the generated summary. Specifically, the server feeds the extracted text data into a BERT model to generate the summary. During this process, the model undergoes multiple validations to ensure that the key points are reflected.

[0151] Step 7: Provide a summary

[0152] The server converts the generated summary into JSON format and sends it to the terminal. The input is the generated summary text, and the output is the summary data sent to the terminal in JSON format. Specifically, the server converts the summary into JSON format and provides it to the user through the web application's user interface. The terminal displays the received summary to the user, who can access detailed information as needed.

[0153] (Application Example 1)

[0154] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0155] Traditionally, the collection and summarization of technical literature and intellectual property information was often done manually, requiring considerable time and effort. Furthermore, there was no way for users to instantly access the latest technical information about products within a virtual store, making effective information gathering and provision difficult. This resulted in reduced user convenience and impaired the efficiency of information gathering.

[0156] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0157] In this invention, the server includes means for collecting technical literature and intellectual property information from a database; means for preprocessing the collected data and removing irrelevant information from the text; means for training a generative artificial intelligence model using the preprocessed data; means for receiving search queries from users; means for extracting relevant information from the database based on the received search queries; means for generating a summary of the extracted information using the generative artificial intelligence model; means for providing the generated summary to the user; and means for enabling the user to instantly obtain the latest technical and intellectual property information about products in a virtual store using a smartphone or smart glasses. This makes it possible for users to efficiently collect technical literature and intellectual property information and instantly check the summarized information.

[0158] A "database" is a system for organizing, systematizing, and centrally managing information.

[0159] "Technical literature" refers to documents that describe research results and knowledge related to science and technology.

[0160] "Intellectual property information" refers to information concerning intellectual property rights such as patents, trademarks, and copyrights.

[0161] "Preprocessing" is the process of removing noise and irrelevant information from raw data and converting it into an analyzable format.

[0162] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates text and other data based on training data.

[0163] A "search query" is a keyword or phrase that a user uses to search for specific information.

[0164] A "summary" is a short, concise version of a longer text, extracting the most important information from it.

[0165] A "smartphone" is a portable device that, in addition to telephone functionality, also possesses advanced computer capabilities.

[0166] "Smart glasses" are glasses-type devices that display information in the field of view and have the ability to connect to the internet.

[0167] A "virtual store" is an online shop that exists on the internet, where users can browse and purchase products in a virtual space.

[0168] A "product" is an item or service that is manufactured and supplied to consumers.

[0169] This invention relates to a system that collects technical literature and intellectual property information and provides summarized information using a generative artificial intelligence model. This system enables users, particularly in virtual stores, to instantly obtain the latest technical and intellectual property information about products using a smartphone or smart glasses.

[0170] System Overview

[0171] 1. Collection from databases

[0172] The server connects to technical literature databases and intellectual property databases to collect relevant information. For example, it queries the databases using keywords such as "autonomous driving" and retrieves corresponding data. The retrieved data is stored in a structured format such as JSON.

[0173] 2. Data preprocessing

[0174] The server converts the collected data into text format. It then cleans the data by removing noise (e.g., HTML tags and special characters). Next, it tokenizes the text data and removes irrelevant stop words. The pre-processed data is then stored back in the database.

[0175] 3. Training the AI ​​model

[0176] The server uses preprocessed data to train a generative artificial intelligence model. This model may be one such model, such as BERT or GPT. Once trained, the model learns important patterns and is deployed for inference.

[0177] 4. Acceptance of user queries

[0178] The terminal receives search queries from users through an input interface. For example, a user might enter keywords such as "latest information on autonomous driving technology."

[0179] 5. Query analysis and extraction of corresponding data

[0180] The server receives search queries sent from the terminal and analyzes the keywords. Next, it extracts relevant technical literature and intellectual property information from the database based on the query.

[0181] 6. Summary Generation

[0182] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The generated summary is then verified to ensure that the main points are reflected before being provided to the user.

[0183] 7. Providing a summary

[0184] The server sends the summarized results, converted to JSON format, to the terminal. The terminal then displays the generated summarized information to the user.

[0185] Specific example

[0186] For example, if the user wants to obtain information about "autonomous driving," the server sends queries to technical literature databases and intellectual property databases, collects papers and patents related to autonomous driving, and generates clean data. A trained generative artificial intelligence model generates a summary from this data, such as "the latest autonomous driving technologies include AI-based real-time road surface detection and high-precision obstacle detection using laser sensors." Users can instantly check this summary information via their smartphone or smart glasses. This invention enables rapid information provision within virtual stores.

[0187] Example of a prompt

[0188] The following prompts are used as input to the generative artificial intelligence model:

[0189] Summarize the latest information on autonomous driving technology.

[0190] The above describes the embodiments of the present invention. This system allows users to efficiently collect and verify the latest technical and intellectual property information within a virtual store.

[0191] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0192] Step 1:

[0193] The server connects to technical literature databases and intellectual property databases to collect relevant information. The server queries the databases using keywords (e.g., "autonomous driving") and stores the retrieved data in JSON format.

[0194] Input: Keyword query

[0195] Output: Technical document data and intellectual property information in JSON format

[0196] Specific operation: The server accesses the database via an API and retrieves and saves the relevant data.

[0197] Step 2:

[0198] The server preprocesses the collected data. Specifically, it converts it to text data, removes noise (e.g., HTML tags and special characters), and removes irrelevant stop words. It also performs tokenization. The preprocessed data is then stored again in the database.

[0199] Input: Technical document data and intellectual property information in JSON format

[0200] Output: Preprocessed text data

[0201] Specific operation: The server performs text data cleanup, tokenization, and stop word removal.

[0202] Step 3:

[0203] The server trains generative artificial intelligence models using preprocessed data. Specifically, it inputs data into models such as BERT and GPT to learn patterns. Once the training is complete, the models are stored on the server and deployed for inference.

[0204] Input: Preprocessed text data

[0205] Output: Trained generative artificial intelligence model

[0206] Specific operation: The server inputs data into the model and executes and completes the training process.

[0207] Step 4:

[0208] The terminal receives search queries from the user. The user enters search keywords (e.g., "latest information on autonomous driving technology") through the interface.

[0209] Input: Search keyword (user query)

[0210] Output: Search query (sent from terminal to server)

[0211] Specific operation: The terminal receives queries through the user interface and sends them to the server.

[0212] Step 5:

[0213] The server receives search queries sent from terminals and analyzes the keywords. Based on the queries, the server extracts relevant technical literature and intellectual property information from its database.

[0214] Input: Search query from the user

[0215] Output: Extracted technical literature and intellectual property information

[0216] Specific operation: The server parses the query and quickly extracts relevant information from the database.

[0217] Step 6:

[0218] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The generated summary is then verified to ensure that the main points are reflected before being provided to the user.

[0219] Input: Extracted technical literature and intellectual property information

[0220] Output: Summarized information

[0221] Specific operation: The server uses a trained generative artificial intelligence model to generate and verify a summary of the extracted data.

[0222] Step 7:

[0223] The server converts the generated summary into JSON format and sends it to the terminal. The terminal then displays the summary information to the user.

[0224] Input: Summary information (generated by a generative AI model)

[0225] Output: Summary information in JSON format (sent from server to terminal)

[0226] Specific operation: The server encodes the summary information and sends it to the terminal. The terminal displays the summary information to the user.

[0227] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0228] This invention relates to a system for collecting technical literature and intellectual property information, enabling users to efficiently obtain this information. The system is characterized by collecting information from a database, performing preprocessing, generating summaries using a generative artificial intelligence model, and providing them to the user. Furthermore, it incorporates an emotion engine that analyzes the user's emotions, and adjusts and customizes the information based on the user's emotional state.

[0229] System programming and processing

[0230] 1. Data Collection

[0231] The server connects to technical literature databases and intellectual property databases to collect relevant information. For example, it uses APIs to retrieve "technical literature" and "intellectual property information" from the databases.

[0232] 2. Data preprocessing

[0233] The server converts the collected data into text format and removes noise (e.g., HTML tags and special characters). Next, it tokenizes the clean data and removes irrelevant stop words.

[0234] 3. Training the AI ​​model

[0235] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model, and the learning process proceeds.

[0236] 4. Acceptance of user queries

[0237] The terminal receives search queries from the user through its interface. For example, the user might enter the query, "Latest trends in autonomous driving technology."

[0238] 5. Query analysis and extraction of corresponding data

[0239] The server analyzes the search query received from the terminal and extracts keywords. Next, it extracts relevant technical literature and intellectual property information from the database based on the query.

[0240] 6. Summary Generation

[0241] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The summary includes key technical points, findings, and application examples.

[0242] 7. Emotion analysis

[0243] The device collects emotional data from the user's interface operations and input.

[0244] The server uses an emotion engine to analyze the user's emotions. For example, it uses text analysis to determine whether the user is experiencing emotions such as "anxiety" or "interest."

[0245] 8. Customization of information provision

[0246] The server customizes the generated summary information based on the analyzed emotional state. For example, if the user is feeling "anxious," it adds reassuring phrases.

[0247] 9. Providing a summary

[0248] The server sends a customized summary to the terminal.

[0249] The device displays summary information to the user.

[0250] Specific example

[0251] Data Acquisition and Preprocessing

[0252] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to autonomous driving. The collected data is stored in JSON format, and after noise reduction and tokenization, it is stored in the database.

[0253] Model training and query analysis

[0254] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize important information related to autonomous driving technology. When a user enters "latest information on autonomous driving technology" as a query, the server parses this query and extracts relevant papers and patents.

[0255] Summarization generation and sentiment analysis

[0256] The extracted information is summarized by a generative artificial intelligence model. Simultaneously, the terminal collects the user's emotions when they enter queries, and the server analyzes those emotions using an emotion engine.

[0257] Customization and delivery

[0258] The summarized information is customized based on the user's emotional state. For example, if the user indicates "anxiety," the summary will include expressions that emphasize stability and reliability. Finally, the customized summary information is sent from the server to the terminal and displayed to the user.

[0259] In summary, the present invention is a system that enables users to effectively collect and summarize technical literature and intellectual property information, as well as to provide customized information based on the user's emotional state.

[0260] The following describes the processing flow.

[0261] Step 1:

[0262] The server connects to technical literature databases and intellectual property databases, and uses APIs and queries to collect "technical literature" and "intellectual property information." For example, it retrieves data related to "autonomous driving."

[0263] Step 2:

[0264] The server converts the collected data into JSON format and stores it securely in local or cloud storage.

[0265] Step 3:

[0266] The server preprocesses the collected data. Specifically, it removes HTML tags and special characters from the text. Furthermore, it performs natural language processing such as tokenization and stop word removal to clean the data.

[0267] Step 4:

[0268] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model to learn important patterns and features.

[0269] Step 5:

[0270] The server evaluates the trained generative artificial intelligence model on a test dataset to verify its accuracy. Once the evaluation is complete, the model is deployed for production use.

[0271] Step 6:

[0272] The terminal receives search queries from the user through its interface. For example, the user enters the query "latest trends in autonomous driving technology."

[0273] Step 7:

[0274] The terminal collects emotional data from the user's facial expressions, typing speed, and context when they enter queries. The collected emotional data is sent to the server.

[0275] Step 8:

[0276] The server analyzes the search queries received from the terminal and extracts "autonomous driving technology" and "latest trends" as keywords. Next, it extracts technical literature and intellectual property information related to these keywords from the database.

[0277] Step 9:

[0278] The server inputs the extracted information into a generative artificial intelligence model to generate a summary. The summary includes the main technical points, discoveries, and application examples.

[0279] Step 10:

[0280] The server analyzes the sentiment data transmitted using a sentiment engine. For example, it may be determined that the user's sentiment is "uneasy".

[0281] Step 11:

[0282] The server customizes the summary generated based on the result of the sentiment analysis. For example, if "uneasy" is identified, words that give a sense of reassurance are added to the summary content.

[0283] Step 12:

[0284] The server converts the customized summary into JSON format and transmits it to the terminal.

[0285] Step 13:

[0286] The terminal displays the received customized summary result on the user interface. The user can obtain information according to their emotional state while checking the summary information for the query.

[0287] Through this series of processes, the user can efficiently collect the latest technical literature and intellectual property information and use the summarized information with confidence. This system realizes a more personalized information provision by considering the user's emotional state.

[0288] (Example 2)

[0289] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart device 14 is referred to as the "terminal".

[0290] In conventional technical information gathering systems, it was difficult to efficiently collect relevant information from vast databases and summarize that information. Furthermore, the summarized information did not adapt to the user's emotional state, resulting in a decline in the quality of information provided. As a result, users were unable to obtain the necessary information quickly and appropriately, leading to wasted effort and time.

[0291] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0292] In this invention, the server includes means for collecting technical documents and intellectual property information from a database, means for preprocessing the collected data and removing irrelevant information from the text, and means for training a generative artificial intelligence model using the preprocessed data. This makes it possible to efficiently preprocess the collected data and train the generative artificial intelligence model with high accuracy. Furthermore, by including means for extracting relevant information from the database based on a received search query and generating a summary of that information using the generative artificial intelligence model, means for analyzing the user's emotional state during summary generation and customizing the summary information based on that state, and means for providing the generated summary to the user, it becomes possible to provide high-quality information adapted to the user's emotions, thereby improving the efficiency and accuracy of information collection.

[0293] A "database" is an information system that systematically stores multiple pieces of data, making it possible to search for and retrieve them.

[0294] "Technical literature" refers to documents such as academic papers, research reports, and technical reports related to a specific technical field.

[0295] "Intellectual property information" refers to data and documents related to intellectual property rights, such as patents, trademarks, and copyrights.

[0296] "Preprocessing" refers to a series of operations that transform data into a format suitable for analysis and model training.

[0297] A "generative artificial intelligence model" is an artificial intelligence model that can generate new text or information based on input data.

[0298] "Tokenization" is the process of splitting text data into meaningful units such as words or phrases.

[0299] "Stop words" are common words that are excluded as having low importance in text analysis and natural language processing.

[0300] A "search query" is a keyword or phrase entered by a user to search for information.

[0301] "Summarization" refers to extracting and presenting the main points of a document or information in a shortened form.

[0302] "Sentiment analysis" is the process of judging the emotional state of a user from text or behavior.

[0303] "Customization" means changing and adjusting information or services according to specific needs or situations of users.

[0304] A "user interface" is a general term for the screens and operating means through which a user interacts with a system.

[0305] "Noise" refers to data or information that is unnecessary or harmful for analysis.

[0306] "Hugging Face" is an organization that provides open-source libraries and tools for natural language processing.

[0307] The "Transformers library" is an open-source software library used for building and training generative artificial intelligence models.

[0308] ElasticSearch (registered trademark) is a full-text search engine and analytics engine, a system designed for high-speed searching of large amounts of data.

[0309] "spaCy" is a high-performance natural language processing (NLP) library used for text analysis and model building.

[0310] "BeautifulSoup" is a Python library for parsing HTML and XML files and extracting specific information.

[0311] TextBlob is a Python library for easily performing sentiment analysis and classification of text data.

[0312] "VADER" is a Python library specifically designed for text sentiment analysis.

[0313] This invention relates to a system for efficiently collecting, summarizing, and providing technical literature and intellectual property information to users. This system is characterized by collecting information from a database, performing preprocessing, generating summaries using a generative artificial intelligence model, and providing information customized based on the user's sentiment.

[0314] Data collection

[0315] The server connects to technical literature databases and intellectual property databases via APIs to collect relevant information. For example, it collects papers and patent data related to "autonomous driving" from databases such as IEEE Xplore and Google® Patents. The server authenticates using an API key, sends search queries, and retrieves data.

[0316] Specific example: The server sends "self-driving cars" as a search query to the IEEE Xplore API to retrieve a list of papers related to autonomous driving.

[0317] Data preprocessing

[0318] The server converts the collected data into text format and removes noise such as HTML tags and special characters using Python's BeautifulSoup library. Next, the clean text data is tokenized, and irrelevant stop words are removed using Python's NLTK library.

[0319] Specific example: Perform the operation of removing HTML tags from acquired research paper data and extracting only the text portion. In the text "The advancements in self-driving cars include sensors and AI technology.", stop words such as "The" and "in" are removed, and the tokens "advancements, self-driving, cars, include, sensors, AI, technology" are extracted.

[0320] Training an AI model

[0321] The server uses preprocessed data to train generative artificial intelligence models (e.g., BERT and GPT). The Hugging Face Transformers library is used to train the generative AI models. The GPT-3® model is trained using preprocessed text data on autonomous driving technology, aiming for highly accurate summary generation.

[0322] Acceptance of user queries

[0323] The terminal accepts search queries from users through a GUI (Graphical User Interface). For example, a user might enter the query "latest trends in autonomous driving technology" and click the search button.

[0324] Specific example: The user enters "latest trends in autonomous driving technology" and presses the search button.

[0325] Query analysis and extraction of corresponding data

[0326] The server parses the search query received from the terminal and extracts keywords from the query using the Python spaCy library. Next, it searches the database using Elasticsearch and extracts relevant technical documents and patent information.

[0327] Specific example: Extract the keyword "autonomous driving technology" from the query "latest trends in autonomous driving technology," and use Elasticsearch to search for papers and patent information related to this keyword.

[0328] Summary generation

[0329] The server inputs the extracted data into a generative artificial intelligence model to generate a summary that includes key technical points, discoveries, and application examples. For example, it can use a GPT-3 model to quickly summarize information and provide it to the user.

[0330] Specific example: Input extracted technical documents and patent information into GPT-3 to generate a "Summary of the Latest Trends in Autonomous Driving Technology."

[0331] Emotion analysis

[0332] The device collects emotional data from the user's text input and actions. For example, it uses an emotional analysis library (e.g., TextBlob or VADER) to analyze whether the user is experiencing emotions such as "anxiety" or "interest."

[0333] Specific example: If a user enters a keyword like "difficult," the TextBlob is used to analyze the emotion "anxiety."

[0334] Customization of information provision

[0335] The server customizes the summary information generated based on the analyzed emotional state. For example, if the user is feeling "anxious," it adds reassuring phrases and detailed information.

[0336] Specific example: Add a sentence to the summary such as, "This technology has already been safely implemented by major manufacturers and has a proven track record."

[0337] Summary

[0338] The server sends customized summary information to the terminal, which then displays it to the user.

[0339] Specific example: Display a customized summary in the user interface of a browser.

[0340] Through the steps described above, this system can efficiently collect and analyze technical literature and intellectual property information, and provide information tailored to the user's emotions. This allows users to quickly obtain the necessary information and deepen their understanding.

[0341] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0342] Program processing flow

[0343] Step 1: Data Collection

[0344] Input: Search queries to technical literature databases and intellectual property databases

[0345] The server connects to technical literature databases and intellectual property databases via APIs and collects data related to the information requested by the user. In this process, the server authenticates using an API key and sends search queries to retrieve data. For example, it might send "self-driving cars" as a query to the IEEE Xplore API.

[0346] Output: Collected technical literature data and intellectual property data

[0347] Step 2: Data Preprocessing

[0348] Input: Collected technical literature data and intellectual property data

[0349] The server converts the collected data into text format and removes noise such as HTML tags and special characters. The Python BeautifulSoup library is used for this process. Furthermore, the clean text data is tokenized, and irrelevant stop words are removed using the NLTK library. In this way, clean data suitable for analysis is obtained.

[0350] Output: Pre-processed clean data

[0351] Step 3: Training the AI ​​model

[0352] Input: Pre-processed clean data

[0353] The server uses preprocessed data to train generative artificial intelligence models (e.g., BERT or GPT). This training utilizes the Hugging Face Transformers library. The generative AI models learn to generate highly accurate summaries based on a large amount of technical literature data.

[0354] Output: Trained generative AI model

[0355] Step 4: Accepting User Queries

[0356] Input: User's search query

[0357] The terminal accepts search queries from users via a GUI. The user enters a query (e.g., "latest trends in autonomous driving technology") into the text input field and clicks the search button.

[0358] Output: User's search query

[0359] Step 5: Query analysis and extraction of corresponding data

[0360] Input: User's search query

[0361] The server parses the search query received from the terminal and extracts keywords from the query using the Python spaCy library. Next, it searches the database using Elasticsearch to extract relevant technical documents and patent information. For example, it extracts the keyword "autonomous driving technology" from the query "latest trends in autonomous driving technology" and performs a search.

[0362] Output: Extracted technical literature and patent information

[0363] Step 6: Summary Generation

[0364] Input: Extracted technical documents and patent information

[0365] The server inputs the extracted data into a generative artificial intelligence model to generate a summary that includes key technical points, discoveries, and application examples. For example, it uses a GPT-3 model to quickly summarize information and provide it to the user.

[0366] Output: Summarized information

[0367] Step 7: Emotion Analysis

[0368] Input: User query input and operation information

[0369] The device collects emotional data from the user's text input and actions. For example, it uses an emotional analysis library (such as TextBlob or VADER) to determine if the user is experiencing emotions like "anxiety" or "interest." If the user types "difficult," the device analyzes that emotion.

[0370] Output: User's emotional state

[0371] Step 8: Customizing the information provided

[0372] Input: Summarized information and user's emotional state

[0373] The server customizes the summary information generated based on the analyzed emotional state. If the user is feeling "anxious," reassuring phrases and detailed information are added to the summary. For example, a sentence like, "This technology has already been securely implemented by major manufacturers and has a proven track record," might be added to the summary.

[0374] Output: Customized summary information

[0375] Step 9: Provide a summary

[0376] Input: Customized summary information

[0377] The server sends customized summary information to the terminal. The terminal displays the received summary information in its user interface, providing the information to the user.

[0378] Output: Summary information displayed to the user

[0379] Through the steps outlined above, this system enables the effective collection, analysis, and summarization of technical literature and intellectual property information, as well as the provision of information based on user sentiment. This allows users to quickly obtain the necessary information and make appropriate decisions.

[0380] (Application Example 2)

[0381] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0382] In today's world, users are required to acquire advanced information quickly and efficiently. However, a vast amount of technical literature and intellectual property information exists, and extracting and summarizing highly relevant information from it requires considerable time and effort. Furthermore, the information provided may lack effectiveness for the recipient because it is not appropriately customized according to the user's emotions and circumstances. In addition, advertising also faces the challenge of not providing appropriate information that is tailored to the user's situation.

[0383] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting technical literature and intellectual property information from a database, means for preprocessing the collected data and removing irrelevant information from the text, means for training a generative artificial intelligence model using the preprocessed data, means for customizing a summary generated based on the user's emotional state, and means for providing the user with advertising information including the customized summary. As a result, the user can efficiently obtain optimized technical literature and intellectual property information, and the receptivity of the information is improved because the most appropriate advertising information is provided according to the user's emotions and situation.

[0384] A "database" is a collection of information, including technical documents and intellectual property information. The information is stored in a searchable format.

[0385] "Technical documents" are materials that describe research results and technical explanations related to science and technology. This includes academic papers and technical reports.

[0386] "Intellectual property information" refers to information concerning intellectual property rights such as patents, trademarks, and copyrights. This includes patent application documents and registration certificates.

[0387] "Preprocessing" is the process of removing irrelevant information and noise from collected data and converting the data into an analyzable format. This also includes the process of removing HTML tags and special characters from text.

[0388] A "generative artificial intelligence model" is a machine learning model that has the ability to generate new information based on large amounts of data. Examples include BERT and GPT.

[0389] A "search query" is a request for input from a user to obtain specific information. It consists of keywords or questions.

[0390] "Summary generation" is the process of shortening long texts or data and extracting the main points. This is done automatically using artificial intelligence models.

[0391] "Emotional state" refers to the psychological state a user is in when receiving information. Specifically, it includes "excitement," "anxiety," and "interest."

[0392] "Customization" refers to adjusting and modifying information to suit the specific needs and feelings of the user.

[0393] "Advertising information" refers to information used to promote products and services. It is provided in an optimized format based on the user's interests and emotional state.

[0394] This invention relates to a system that collects technical literature and intellectual property information from a database and provides information optimized for the user. This system preprocesses the collected data, generates summaries using a generative artificial intelligence model, further customizes the information based on the user's emotional state, and provides advertising information along with it.

[0395] Hardware and software usage

[0396] The system is implemented using the following hardware and software:

[0397] Server: Responsible for large-scale data collection, preprocessing, training of generative artificial intelligence models, and summary generation.

[0398] Terminal: Responsible for receiving search queries from users and collecting sentiment data.

[0399] Generative artificial intelligence models: For example, BERT or GPT are used.

[0400] Emotion analysis engine: Emotion analysis is performed using the Hugging Face Transformers library.

[0401] Database: Stores technical literature and intellectual property information, and provides data via API.

[0402] Data processing and data calculation

[0403] Data Acquisition and Preprocessing

[0404] The server collects technical literature and intellectual property information from a database via an API. The collected data is then stored in JavaScript Object Notation (JSON) format and converted to text format. The text undergoes preprocessing such as noise reduction, tokenization, and stop word removal.

[0405] Training of generative artificial intelligence models

[0406] Using preprocessed data, the server trains a generative artificial intelligence model. This training is performed to generate summaries related to a specific technical domain. For example, PyTorch is used to train the model.

[0407] Emotion analysis

[0408] The device collects emotional data in real time as the user enters search queries. This emotional data is input into an emotion analysis engine through text analysis. The emotion engine detects what emotional state the user is in, such as "excitement," "interest," or "anxiety."

[0409] Customized information provision

[0410] The server customizes the generated summary information based on sentiment analysis results. For example, it adds reassuring or attention-grabbing phrases to the summary. It then sends the customized summary along with advertising information to the device.

[0411] Specific example

[0412] When a user enters the query "latest smartphone technology," the server collects relevant information from technical literature and intellectual property databases. The collected data is then preprocessed and summarized by a generative artificial intelligence model. After the summary is generated, an emotion analysis engine analyzes the user's emotional state, and the summary information is customized based on the results.

[0413] For example, if a user expresses excitement, the summary information will include phrases that emphasize the appeal and future potential of the new technology. Finally, the customized summary information and related advertisements are sent to the device and displayed to the user.

[0414] Example of a prompt:

[0415] "Could you tell me about the latest trends in smartphone technology?"

[0416] The above describes the specific forms for carrying out the invention.

[0417] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0418] Step 1: Data Collection

[0419] The server collects relevant information from technical literature databases and intellectual property information databases via APIs. The input is a query containing keywords related to the technical literature and intellectual property information to be collected. The server stores the data collected via the API in JSON format. The output is a dataset of the collected technical literature and intellectual property information.

[0420] Step 2: Data Preprocessing

[0421] The server preprocesses the data collected in Step 1. Specifically, it removes noise such as HTML tags and special characters from the text and performs tokenization. The collected dataset of technical documents and intellectual property information is used as input. The output is clean data that has been de-noised and tokenized.

[0422] Step 3: Training the Generative AI Model

[0423] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using preprocessed data. The input is a preprocessed dataset. Training is performed using machine learning libraries such as PyTorch. The output is a generative artificial intelligence model trained to generate summaries corresponding to a specific technology domain.

[0424] Step 4: Accepting User Queries

[0425] The terminal receives search queries from the user through its interface. The input is the search query entered by the user (e.g., "Please tell me about the latest trends in smartphone technology."). The output is the received search query.

[0426] Step 5: Query analysis and data extraction

[0427] The server parses the search query received from the terminal and extracts keywords. It uses the user's search query as input. Next, the server extracts relevant technical literature and intellectual property information from the database based on the query. The output is a set of relevant technical literature and intellectual property information.

[0428] Step 6: Summary Generation

[0429] The server inputs extracted technical literature and intellectual property information into a generative artificial intelligence model to generate a summary. The extracted dataset and the trained generative AI model are used as input. The output is a summary containing the key points.

[0430] Step 7: Emotion Analysis

[0431] The terminal collects sentiment data when the user enters a search query. The input consists of user interface interactions and input text. The server inputs the collected sentiment data into a sentiment analysis engine to analyze the user's emotional state (e.g., "excited," "anxious"). The output is the user's emotional state.

[0432] Step 8: Customizing the information provided

[0433] The server customizes the generated summary information based on the analyzed user's emotional state. It uses the generated summary and the user's emotional state as input. For example, if the user is "excited," it adds more appealing expressions to the summary. The output is the customized summary information.

[0434] Step 9: Provide a summary

[0435] The server sends customized summary information and associated advertising information to the device. The device uses the customized summary information and generated advertising information as input. The device displays this information to the user. The output is the customized summary information and advertising information provided to the user.

[0436] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0437] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0438] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0439] [Second Embodiment]

[0440] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0441] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0442] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0443] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0444] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0445] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0446] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0447] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0448] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0449] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0450] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0451] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0452] This invention relates to a system for collecting technical documents and intellectual property information, enabling users to efficiently obtain this information. The system is characterized by collecting information from a database, performing preprocessing, generating a summary using a generative artificial intelligence model, and providing it to the user.

[0453] System Programs and Processing

[0454] 1. Data Collection

[0455] The server connects to technical literature databases and intellectual property databases to collect relevant information.

[0456] The server queries the database using keywords such as "autonomous driving" and retrieves the corresponding data.

[0457] The server stores the retrieved data in a structured format such as JSON.

[0458] 2. Data preprocessing

[0459] The server converts the collected data into text format.

[0460] The server cleans up the data by removing noise from the text (e.g., HTML tags and special characters).

[0461] The server tokenizes the clean data and removes irrelevant stop words.

[0462] The server saves the pre-processed data to the database.

[0463] 3. Training the AI ​​model

[0464] The server uses the pre-processed dataset to train a generative artificial intelligence model.

[0465] The server uses generative artificial intelligence models (e.g., BERT or GPT) to learn important patterns in the data.

[0466] The server stores the trained model and deploys it for inference.

[0467] 4. Acceptance of user queries

[0468] The terminal receives search queries from the user through the input interface.

[0469] The user enters keywords such as "latest information on autonomous driving technology."

[0470] 5. Query analysis and extraction of corresponding data

[0471] The server receives the search query sent from the terminal and parses the keywords.

[0472] The server extracts relevant technical literature and intellectual property information from the database based on the query.

[0473] 6. Summary Generation

[0474] The server inputs the extracted data into a generative artificial intelligence model and generates a summary.

[0475] The server verifies the summary results to ensure that the main points are reflected.

[0476] 7. Providing a summary

[0477] The server converts the generated summary into JSON format and sends it to the terminal.

[0478] The device displays summary information to the user.

[0479] Specific example

[0480] Data Acquisition and Preprocessing

[0481] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to autonomous driving. The collected data is stored in JSON format, and after noise reduction and tokenization, it is stored in the database.

[0482] Model training and query analysis

[0483] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize important information related to autonomous driving technology. When a user enters "latest information on autonomous driving technology" as a query, the server analyzes this query and extracts relevant papers and patents.

[0484] Summary generation and delivery

[0485] The extracted information is summarized by a generative artificial intelligence model. For example, a summary might be generated stating, "The latest autonomous driving technologies include AI-powered real-time road surface detection and high-precision obstacle detection using laser sensors." The generated summary is sent from the server to the terminal, where the user can view it through the interface.

[0486] Based on the above, the present invention is a system that enables users to effectively collect and summarize technical documents and intellectual property information.

[0487] The following describes the processing flow.

[0488] Step 1:

[0489] The server connects to technical literature databases and intellectual property databases. For example, it uses APIs to collect data from the "technical literature database" and the "intellectual property database."

[0490] Step 2:

[0491] The server converts the collected data into JSON format and saves it to local or cloud storage. The collected data includes the title of the paper, authors, abstracts, patent numbers, and summaries of the inventions.

[0492] Step 3:

[0493] The server preprocesses the collected data. Specifically, it removes HTML tags and special characters from the text. Next, it tokenizes the cleansed data and removes unwanted words (stop words).

[0494] Step 4:

[0495] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model to begin the learning process.

[0496] Step 5:

[0497] The server evaluates the trained generative artificial intelligence model on a test dataset to verify its accuracy. Once the evaluation is complete, the model is deployed for inference.

[0498] Step 6:

[0499] The terminal receives search queries from users through its interface. For example, it might accept a query such as "latest trends in autonomous driving technology."

[0500] Step 7:

[0501] The server analyzes the search query received from the terminal and extracts keywords. For example, "autonomous driving technology" and "latest trends" might be extracted as keywords.

[0502] Step 8:

[0503] The server extracts relevant technical literature and intellectual property information from the database based on the analyzed query. The extracted data is temporarily loaded into memory.

[0504] Step 9:

[0505] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The summary includes key technical points, findings, and application examples.

[0506] Step 10:

[0507] The server converts the generated summary into JSON format and sends it to the terminal. Because the transmitted summary data is structured, it is displayed in a user-friendly format.

[0508] Step 11:

[0509] The terminal displays the received summary results in the user interface. Users can review the summary information corresponding to their search queries and quickly obtain the necessary technical information.

[0510] This series of processes allows users to efficiently collect the latest technical literature and intellectual property information, and obtain the necessary information in a summarized format.

[0511] (Example 1)

[0512] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0513] Because technical literature and intellectual property information are vast, it is difficult for individual users to efficiently collect the necessary information and summarize and understand the key points. Traditional systems primarily rely on manual information gathering and summarization, which is time-consuming and labor-intensive. Furthermore, the process of extracting and summarizing appropriate information based on search queries is often inefficient. Therefore, there is a need for a system that can automatically collect, preprocess, and summarize technical literature and intellectual property information, and efficiently provide it to users.

[0514] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0515] In this invention, the server includes means for collecting relevant information from a technical literature database and an intellectual property database; means for converting the collected information into text format, removing noise, and creating clean data; and means for tokenizing the clean data and preprocessing irrelevant information. This makes it possible to automatically collect and preprocess technical literature and intellectual property information to generate a dataset.

[0516] The server also includes means for training a generative artificial intelligence model using a preprocessed dataset, means for receiving search queries from users, means for analyzing the received search queries and extracting relevant information from technical literature databases and intellectual property databases, means for generating summaries of the extracted information using the generative artificial intelligence model, and means for converting the generated summaries into JSON format and providing them to users. This makes it possible to efficiently extract relevant information, summarize the key points, and provide them to users.

[0517] A "technical literature database" is a database that collects technical information, such as academic papers and technical reports, and stores it in a searchable format.

[0518] An "intellectual property database" is a database that collects and stores information related to intellectual property, such as patents, trademarks, and copyrights, in a searchable format.

[0519] A "generative artificial intelligence model" is an artificial intelligence model that has the ability to generate new data in tasks such as natural language processing and image generation.

[0520] "JSON format" is an abbreviation for JavaScript Object Notation, and it is a lightweight data exchange format for structuring and storing data.

[0521] "Noise reduction" is the process of removing unnecessary information (e.g., HTML tags and special characters) from data to generate clean data.

[0522] A "query" is a type of inquiry that a user enters in a database or search system to retrieve specific information.

[0523] "Tokenization" is the process of dividing text data into semantic units (tokens) such as words and phrases.

[0524] "Preprocessing" refers to the process of preparing data, such as cleaning up and standardizing its format, before analyzing the data or training models.

[0525] A "summary" is information that extracts the important points from a vast amount of information and presents them in a short format.

[0526] "Analysis" is the process of analyzing data and query content to understand and process their meaning and structure.

[0527] "Extraction" is the process of retrieving data or information based on specific conditions.

[0528] This invention relates to a system for efficiently collecting and processing technical literature and intellectual property information and providing it to users in a summarized format. The system collects relevant information from technical literature databases and intellectual property databases, performs preprocessing, generates summaries using a generative artificial intelligence model, and provides them to users.

[0529] Data collection

[0530] The server periodically accesses technical literature databases (e.g., IEEE Xplore) and intellectual property databases (e.g., USPTO) to collect relevant information. For example, the server automatically sends queries at 3:00 AM every day to retrieve the latest papers and patent information related to keywords such as "autonomous driving" and "AI technology." The collected information is temporarily stored in data storage in a structured JSON format.

[0531] Data preprocessing

[0532] The server converts the collected data into text format and removes noise. First, HTML tags and special characters are removed using regular expressions. Next, the text data is tokenized (divided into words and phrases), and irrelevant stop words (e.g., "a" and "the" in English) are removed using the NLTK library. The pre-processed data is then saved back to the database.

[0533] Training an AI model

[0534] The server trains a generative artificial intelligence (AI) model using a pre-processed dataset. For example, it trains a BERT model using the Hugging Face Transformers library. The server uses 80% of the data for training and 20% for validation. Once the training is complete, the model is stored in the Hugging Face model hub.

[0535] Acceptance of user queries

[0536] The terminal receives search queries from the user through an input interface. For example, it might be implemented as a web application running in a browser, where the user enters a query such as "latest autonomous driving technology" into the search bar. This query is sent to the server in real time.

[0537] Query analysis and extraction of corresponding data

[0538] The server receives search queries sent from the terminal and analyzes the keywords using TF-IDF (Term Frequency-Inverse Document Frequency). Based on the analysis results, the server extracts relevant information from the database. For example, it might pick out literature and patents related to "latest autonomous driving technology."

[0539] Summary generation

[0540] The server inputs the extracted literature and patent information into a generative artificial intelligence model to generate a summary. Specifically, it feeds the extracted text into a BERT model to generate a summarized text. During this process, the model undergoes multiple validations to ensure that all important points are included.

[0541] Summary

[0542] The server converts the generated summary into JSON format and sends it to the terminal. The terminal displays this summary information to the user. For example, in the user interface of a web application, the summary text is displayed in an easy-to-read format. The user can review the displayed summary and access detailed information as needed.

[0543] Specific example

[0544] Data Acquisition and Preprocessing

[0545] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to "autonomous driving." The collected data is stored in JSON format and, after preprocessing such as removing HTML tags and special characters and tokenizing the text, is stored in the database.

[0546] Model training and query analysis

[0547] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize key information related to "autonomous driving technology." When a user enters "latest information on autonomous driving technology" as a query, the server analyzes this query and extracts relevant papers and patents.

[0548] Summary generation and delivery

[0549] The extracted information is summarized by a generative artificial intelligence model. For example, a summary might be generated stating, "The latest autonomous driving technologies include AI-powered real-time road surface detection and high-precision obstacle detection using laser sensors." The generated summary is sent from the server to the terminal, where the user can view it in a web browser.

[0550] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0551] Step 1: Data collection from the database

[0552] The server connects to technical literature databases and intellectual property databases to collect information related to specific keywords (e.g., "autonomous driving" or "AI technology"). The input is a keyword query, and the output is the collected data in JSON format. Specifically, the server automatically sends queries at 3 AM every day and temporarily stores the retrieved papers and patent information in data storage.

[0553] Step 2: Data Preprocessing

[0554] The server converts the collected data into text format and removes noise. The input is collected data in JSON format, and the output is clean text data. Specifically, the server uses regular expressions to remove HTML tags and special characters, then tokenizes the text data into words and phrases, and removes irrelevant stop words using the NLTK library. This clean data is then stored again in the database.

[0555] Step 3: Training the AI ​​model

[0556] The server trains a generative artificial intelligence model using a preprocessed dataset. The input is a preprocessed dataset, and the output is a trained generative AI model. Specifically, the server trains a BERT model using the Hugging Face Transformers library, using 80% of the data for training and 20% for validation. The trained model is stored in the Hugging Face model hub.

[0557] Step 4: Accepting User Queries

[0558] The terminal receives search queries from the user through an input interface. The input is the user's search query (e.g., "latest autonomous driving technology"), and the output is that query being sent to the server. Specifically, the user enters the query into the search bar of the web application, and that query is sent to the server in real time.

[0559] Step 5: Query analysis and extraction of corresponding data

[0560] The server receives search queries sent from terminals and parses the keywords. The input is the user's search query, and the output is relevant information extracted from the database. Specifically, the server uses TF-IDF (Term Frequency-Inverse Document Frequency) to parse the query and extracts relevant technical documents and patent information from the database based on the results.

[0561] Step 6: Summary Generation

[0562] The server inputs the extracted information into a generative artificial intelligence model to generate a summary. The input includes extracted technical documents and patent information, and the output is the generated summary. Specifically, the server feeds the extracted text data into a BERT model to generate the summary. During this process, the model undergoes multiple validations to ensure that the key points are reflected.

[0563] Step 7: Provide a summary

[0564] The server converts the generated summary into JSON format and sends it to the terminal. The input is the generated summary text, and the output is the summary data sent to the terminal in JSON format. Specifically, the server converts the summary into JSON format and provides it to the user through the web application's user interface. The terminal displays the received summary to the user, who can access detailed information as needed.

[0565] (Application Example 1)

[0566] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0567] Traditionally, the collection and summarization of technical literature and intellectual property information was often done manually, requiring considerable time and effort. Furthermore, there was no way for users to instantly access the latest technical information about products within a virtual store, making effective information gathering and provision difficult. This resulted in reduced user convenience and impaired the efficiency of information gathering.

[0568] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0569] In this invention, the server includes means for collecting technical literature and intellectual property information from a database; means for preprocessing the collected data and removing irrelevant information from the text; means for training a generative artificial intelligence model using the preprocessed data; means for receiving search queries from users; means for extracting relevant information from the database based on the received search queries; means for generating a summary of the extracted information using the generative artificial intelligence model; means for providing the generated summary to the user; and means for enabling the user to instantly obtain the latest technical and intellectual property information about products in a virtual store using a smartphone or smart glasses. This makes it possible for users to efficiently collect technical literature and intellectual property information and instantly check the summarized information.

[0570] A "database" is a system for organizing, systematizing, and centrally managing information.

[0571] "Technical literature" refers to documents that describe research results and knowledge related to science and technology.

[0572] "Intellectual property information" refers to information concerning intellectual property rights such as patents, trademarks, and copyrights.

[0573] "Preprocessing" is the process of removing noise and irrelevant information from raw data and converting it into an analyzable format.

[0574] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates text and other data based on training data.

[0575] A "search query" is a keyword or phrase that a user uses to search for specific information.

[0576] A "summary" is a short, concise version of a longer text, extracting the most important information from it.

[0577] A "smartphone" is a portable device that, in addition to telephone functionality, also possesses advanced computer capabilities.

[0578] "Smart glasses" are glasses-type devices that display information in the field of view and have the ability to connect to the internet.

[0579] A "virtual store" is an online shop that exists on the internet, where users can browse and purchase products in a virtual space.

[0580] A "product" is an item or service that is manufactured and supplied to consumers.

[0581] This invention relates to a system that collects technical literature and intellectual property information and provides summarized information using a generative artificial intelligence model. This system enables users, particularly in virtual stores, to instantly obtain the latest technical and intellectual property information about products using a smartphone or smart glasses.

[0582] System Overview

[0583] 1. Collection from databases

[0584] The server connects to technical literature databases and intellectual property databases to collect relevant information. For example, it queries the databases using keywords such as "autonomous driving" and retrieves corresponding data. The retrieved data is stored in a structured format such as JSON.

[0585] 2. Data preprocessing

[0586] The server converts the collected data into text format. It then cleans the data by removing noise (e.g., HTML tags and special characters). Next, it tokenizes the text data and removes irrelevant stop words. The pre-processed data is then stored back in the database.

[0587] 3. Training the AI ​​model

[0588] The server uses preprocessed data to train a generative artificial intelligence model. This model may be one such model, such as BERT or GPT. Once trained, the model learns important patterns and is deployed for inference.

[0589] 4. Acceptance of user queries

[0590] The terminal receives search queries from users through an input interface. For example, a user might enter keywords such as "latest information on autonomous driving technology."

[0591] 5. Query analysis and extraction of corresponding data

[0592] The server receives search queries sent from the terminal and analyzes the keywords. Next, it extracts relevant technical literature and intellectual property information from the database based on the query.

[0593] 6. Summary Generation

[0594] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The generated summary is then verified to ensure that the main points are reflected before being provided to the user.

[0595] 7. Providing a summary

[0596] The server sends the summarized results, converted to JSON format, to the terminal. The terminal then displays the generated summarized information to the user.

[0597] Specific example

[0598] For example, if the user wants to obtain information about "autonomous driving," the server sends queries to technical literature databases and intellectual property databases, collects papers and patents related to autonomous driving, and generates clean data. A trained generative artificial intelligence model generates a summary from this data, such as "the latest autonomous driving technologies include AI-based real-time road surface detection and high-precision obstacle detection using laser sensors." Users can instantly check this summary information via their smartphone or smart glasses. This invention enables rapid information provision within virtual stores.

[0599] Example of a prompt

[0600] The following prompts are used as input to the generative artificial intelligence model:

[0601] Summarize the latest information on autonomous driving technology.

[0602] The above describes the embodiments of the present invention. This system allows users to efficiently collect and verify the latest technical and intellectual property information within a virtual store.

[0603] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0604] Step 1:

[0605] The server connects to technical literature databases and intellectual property databases to collect relevant information. The server queries the databases using keywords (e.g., "autonomous driving") and stores the retrieved data in JSON format.

[0606] Input: Keyword query

[0607] Output: Technical document data and intellectual property information in JSON format

[0608] Specific operation: The server accesses the database via an API and retrieves and saves the relevant data.

[0609] Step 2:

[0610] The server preprocesses the collected data. Specifically, it converts it to text data, removes noise (e.g., HTML tags and special characters), and removes irrelevant stop words. It also performs tokenization. The preprocessed data is then stored again in the database.

[0611] Input: Technical document data and intellectual property information in JSON format

[0612] Output: Preprocessed text data

[0613] Specific operation: The server performs text data cleanup, tokenization, and stop word removal.

[0614] Step 3:

[0615] The server trains generative artificial intelligence models using preprocessed data. Specifically, it inputs data into models such as BERT and GPT to learn patterns. Once the training is complete, the models are stored on the server and deployed for inference.

[0616] Input: Preprocessed text data

[0617] Output: Trained generative artificial intelligence model

[0618] Specific operation: The server inputs data into the model and executes and completes the training process.

[0619] Step 4:

[0620] The terminal receives search queries from the user. The user enters search keywords (e.g., "latest information on autonomous driving technology") through the interface.

[0621] Input: Search keyword (user query)

[0622] Output: Search query (sent from terminal to server)

[0623] Specific operation: The terminal receives queries through the user interface and sends them to the server.

[0624] Step 5:

[0625] The server receives search queries sent from terminals and analyzes the keywords. Based on the queries, the server extracts relevant technical literature and intellectual property information from its database.

[0626] Input: Search query from the user

[0627] Output: Extracted technical literature and intellectual property information

[0628] Specific operation: The server parses the query and quickly extracts relevant information from the database.

[0629] Step 6:

[0630] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The generated summary is then verified to ensure that the main points are reflected before being provided to the user.

[0631] Input: Extracted technical literature and intellectual property information

[0632] Output: Summarized information

[0633] Specific operation: The server uses a trained generative artificial intelligence model to generate and verify a summary of the extracted data.

[0634] Step 7:

[0635] The server converts the generated summary into JSON format and sends it to the terminal. The terminal then displays the summary information to the user.

[0636] Input: Summary information (generated by a generative AI model)

[0637] Output: Summary information in JSON format (sent from server to terminal)

[0638] Specific operation: The server encodes the summary information and sends it to the terminal. The terminal displays the summary information to the user.

[0639] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0640] This invention relates to a system for collecting technical literature and intellectual property information, enabling users to efficiently obtain this information. The system is characterized by collecting information from a database, performing preprocessing, generating summaries using a generative artificial intelligence model, and providing them to the user. Furthermore, it incorporates an emotion engine that analyzes the user's emotions, and adjusts and customizes the information based on the user's emotional state.

[0641] System Programs and Processing

[0642] 1. Data Collection

[0643] The server connects to technical literature databases and intellectual property databases to collect relevant information. For example, it uses APIs to retrieve "technical literature" and "intellectual property information" from the databases.

[0644] 2. Data preprocessing

[0645] The server converts the collected data into text format and removes noise (e.g., HTML tags and special characters). Next, it tokenizes the clean data and removes irrelevant stop words.

[0646] 3. Training the AI ​​model

[0647] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model, and the learning process proceeds.

[0648] 4. Acceptance of user queries

[0649] The terminal receives search queries from the user through its interface. For example, the user might enter the query, "Latest trends in autonomous driving technology."

[0650] 5. Query analysis and extraction of corresponding data

[0651] The server analyzes the search query received from the terminal and extracts keywords. Next, it extracts relevant technical literature and intellectual property information from the database based on the query.

[0652] 6. Summary Generation

[0653] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The summary includes key technical points, findings, and application examples.

[0654] 7. Emotion analysis

[0655] The device collects emotional data from the user's interface operations and input.

[0656] The server uses an emotion engine to analyze the user's emotions. For example, it uses text analysis to determine whether the user is experiencing emotions such as "anxiety" or "interest."

[0657] 8. Customization of information provision

[0658] The server customizes the generated summary information based on the analyzed emotional state. For example, if the user is feeling "anxious," it adds reassuring phrases.

[0659] 9. Providing a summary

[0660] The server sends a customized summary to the terminal.

[0661] The device displays summary information to the user.

[0662] Specific example

[0663] Data Acquisition and Preprocessing

[0664] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to autonomous driving. The collected data is stored in JSON format, and after noise reduction and tokenization, it is stored in the database.

[0665] Model training and query analysis

[0666] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize important information related to autonomous driving technology. When a user enters "latest information on autonomous driving technology" as a query, the server analyzes this query and extracts relevant papers and patents.

[0667] Summarization generation and sentiment analysis

[0668] The extracted information is summarized by a generative artificial intelligence model. Simultaneously, the terminal collects the user's emotions when they enter queries, and the server analyzes those emotions using an emotion engine.

[0669] Customization and delivery

[0670] The summarized information is customized based on the user's emotional state. For example, if the user indicates "anxiety," the summary will include expressions that emphasize stability and reliability. Finally, the customized summary information is sent from the server to the terminal and displayed to the user.

[0671] In summary, the present invention is a system that enables users to effectively collect and summarize technical literature and intellectual property information, as well as to provide customized information based on the user's emotional state.

[0672] The following describes the processing flow.

[0673] Step 1:

[0674] The server connects to technical literature databases and intellectual property databases, and uses APIs and queries to collect "technical literature" and "intellectual property information." For example, it retrieves data related to "autonomous driving."

[0675] Step 2:

[0676] The server converts the collected data into JSON format and stores it securely in local or cloud storage.

[0677] Step 3:

[0678] The server preprocesses the collected data. Specifically, it removes HTML tags and special characters from the text. Furthermore, it performs natural language processing such as tokenization and stop word removal to clean the data.

[0679] Step 4:

[0680] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model to learn important patterns and features.

[0681] Step 5:

[0682] The server evaluates the trained generative artificial intelligence model on a test dataset to verify its accuracy. Once the evaluation is complete, the model is deployed for production use.

[0683] Step 6:

[0684] The terminal receives search queries from the user through its interface. For example, the user enters the query "latest trends in autonomous driving technology."

[0685] Step 7:

[0686] The terminal collects emotional data from the user's facial expressions, typing speed, and context when they enter queries. The collected emotional data is sent to the server.

[0687] Step 8:

[0688] The server analyzes the search queries received from the terminal and extracts "autonomous driving technology" and "latest trends" as keywords. Next, it extracts technical literature and intellectual property information related to these keywords from the database.

[0689] Step 9:

[0690] The server inputs the extracted information into a generative artificial intelligence model to generate a summary. The summary includes key technical points, findings, and application examples.

[0691] Step 10:

[0692] The server uses an emotion engine to analyze the transmitted emotion data. For example, it might identify the user's emotion as "anxiety."

[0693] Step 11:

[0694] The server customizes the summary generated based on the sentiment analysis results. For example, if "anxiety" is identified, it adds reassuring wording to the summary.

[0695] Step 12:

[0696] The server converts the customized summary into JSON format and sends it to the terminal.

[0697] Step 13:

[0698] The device displays the received, customized summary results in the user interface. Users can review the summary information for their queries and obtain information tailored to their emotional state.

[0699] This series of processes allows users to efficiently collect the latest technical literature and intellectual property information and confidently utilize the summarized information. The system also takes the user's emotional state into account, enabling more personalized information delivery.

[0700] (Example 2)

[0701] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0702] In conventional technical information gathering systems, it was difficult to efficiently collect relevant information from vast databases and summarize that information. Furthermore, the summarized information did not adapt to the user's emotional state, resulting in a decline in the quality of information provided. As a result, users were unable to obtain the necessary information quickly and appropriately, leading to wasted effort and time.

[0703] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0704] In this invention, the server includes means for collecting technical documents and intellectual property information from a database, means for preprocessing the collected data and removing irrelevant information from the text, and means for training a generative artificial intelligence model using the preprocessed data. This makes it possible to efficiently preprocess the collected data and train the generative artificial intelligence model with high accuracy. Furthermore, by including means for extracting relevant information from the database based on a received search query and generating a summary of that information using the generative artificial intelligence model, means for analyzing the user's emotional state during summary generation and customizing the summary information based on that state, and means for providing the generated summary to the user, it becomes possible to provide high-quality information adapted to the user's emotions, thereby improving the efficiency and accuracy of information collection.

[0705] A "database" is an information system that systematically stores multiple pieces of data, making it possible to search for and retrieve them.

[0706] "Technical literature" refers to documents such as academic papers, research reports, and technical reports related to a specific technical field.

[0707] "Intellectual property information" refers to data and documents related to intellectual property rights, such as patents, trademarks, and copyrights.

[0708] "Preprocessing" refers to a series of operations that transform data into a format suitable for analysis and model training.

[0709] A "generative artificial intelligence model" is an artificial intelligence model that can generate new text or information based on input data.

[0710] "Tokenization" is the process of dividing text data into meaningful units such as words or phrases.

[0711] A "stop word" is a common word that is considered low in importance and is excluded in text analysis and natural language processing.

[0712] A "search query" is a keyword or phrase that a user enters to search for information.

[0713] A "summary" is a shortened version of a document or piece of information, containing the main points extracted from it.

[0714] "Sentiment analysis" is the process of determining a user's emotional state from their text and behavior.

[0715] "Customization" refers to modifying or adjusting information and services to suit the specific needs and circumstances of the user.

[0716] "User interface" is a general term for the screens and operating methods that users use to interact with a system.

[0717] "Noise" refers to data or information that is unnecessary or harmful to the analysis.

[0718] Hugging Face is an organization that provides open-source libraries and tools for natural language processing.

[0719] The "Transformers library" is an open-source software library used for building and training generative artificial intelligence models.

[0720] Elasticsearch is a full-text search engine and analytics engine, a system designed for high-speed searching of large amounts of data.

[0721] "spaCy" is a high-performance natural language processing (NLP) library used for text analysis and model building.

[0722] "BeautifulSoup" is a Python library for parsing HTML and XML files and extracting specific information.

[0723] TextBlob is a Python library for easily performing sentiment analysis and classification of text data.

[0724] "VADER" is a Python library specifically designed for text sentiment analysis.

[0725] This invention relates to a system for efficiently collecting, summarizing, and providing technical literature and intellectual property information to users. This system is characterized by collecting information from a database, performing preprocessing, generating summaries using a generative artificial intelligence model, and providing information customized based on the user's sentiment.

[0726] Data collection

[0727] The server connects to technical literature databases and intellectual property databases via APIs to collect relevant information. For example, it collects papers and patent data related to "autonomous driving" from databases such as IEEE Xplore and Google Patents. The server authenticates using an API key, sends search queries, and retrieves data.

[0728] Specific example: The server sends "self-driving cars" as a search query to the IEEE Xplore API to retrieve a list of papers related to autonomous driving.

[0729] Data preprocessing

[0730] The server converts the collected data into text format and removes noise such as HTML tags and special characters using Python's BeautifulSoup library. Next, the clean text data is tokenized, and irrelevant stop words are removed using Python's NLTK library.

[0731] Specific example: Perform the operation of removing HTML tags from acquired research paper data and extracting only the text portion. In the text "The advancements in self-driving cars include sensors and AI technology.", stop words such as "The" and "in" are removed, and the tokens "advancements, self-driving, cars, include, sensors, AI, technology" are extracted.

[0732] Training an AI model

[0733] The server uses preprocessed data to train generative artificial intelligence models (e.g., BERT or GPT). The Hugging Face Transformers library is used to train the generative AI models. The GPT-3 model is trained using preprocessed text data on autonomous driving technology, aiming for highly accurate summary generation.

[0734] Acceptance of user queries

[0735] The terminal accepts search queries from users through a GUI (Graphical User Interface). For example, a user might enter the query "latest trends in autonomous driving technology" and click the search button.

[0736] Specific example: The user enters "latest trends in autonomous driving technology" and presses the search button.

[0737] Query analysis and extraction of corresponding data

[0738] The server parses the search query received from the terminal and extracts keywords from the query using the Python spaCy library. Next, it searches the database using Elasticsearch and extracts relevant technical documents and patent information.

[0739] Specific example: Extract the keyword "autonomous driving technology" from the query "latest trends in autonomous driving technology," and use Elasticsearch to search for papers and patent information related to this keyword.

[0740] Summary generation

[0741] The server inputs the extracted data into a generative artificial intelligence model to generate a summary that includes key technical points, discoveries, and application examples. For example, it can use a GPT-3 model to quickly summarize information and provide it to the user.

[0742] Specific example: Input extracted technical documents and patent information into GPT-3 to generate a "Summary of the Latest Trends in Autonomous Driving Technology."

[0743] Emotion analysis

[0744] The device collects emotional data from the user's text input and actions. For example, it uses an emotional analysis library (e.g., TextBlob or VADER) to analyze whether the user is experiencing emotions such as "anxiety" or "interest."

[0745] Specific example: If a user enters a keyword like "difficult," the TextBlob is used to analyze the emotion "anxiety."

[0746] Customization of information provision

[0747] The server customizes the summary information generated based on the analyzed emotional state. For example, if the user is feeling "anxious," it adds reassuring phrases and detailed information.

[0748] Specific example: Add a sentence to the summary such as, "This technology has already been safely implemented by major manufacturers and has a proven track record."

[0749] Summary

[0750] The server sends customized summary information to the terminal, which then displays it to the user.

[0751] Specific example: Display a customized summary in the user interface of a browser.

[0752] Through the steps described above, this system can efficiently collect and analyze technical literature and intellectual property information, and provide information tailored to the user's emotions. This allows users to quickly obtain the necessary information and deepen their understanding.

[0753] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0754] Program processing flow

[0755] Step 1: Data Collection

[0756] Input: Search queries to technical literature databases and intellectual property databases

[0757] The server connects to technical literature databases and intellectual property databases via APIs and collects data related to the information requested by the user. In this process, the server authenticates using an API key and sends search queries to retrieve data. For example, it might send "self-driving cars" as a query to the IEEE Xplore API.

[0758] Output: Collected technical literature data and intellectual property data

[0759] Step 2: Data Preprocessing

[0760] Input: Collected technical literature data and intellectual property data

[0761] The server converts the collected data into text format and removes noise such as HTML tags and special characters. The Python BeautifulSoup library is used for this process. Furthermore, the clean text data is tokenized, and irrelevant stop words are removed using the NLTK library. In this way, clean data suitable for analysis is obtained.

[0762] Output: Pre-processed clean data

[0763] Step 3: Training the AI ​​model

[0764] Input: Pre-processed clean data

[0765] The server uses preprocessed data to train generative artificial intelligence models (e.g., BERT or GPT). This training utilizes the Hugging Face Transformers library. The generative AI models learn to generate highly accurate summaries based on a large amount of technical literature data.

[0766] Output: Trained generative AI model

[0767] Step 4: Accepting User Queries

[0768] Input: User's search query

[0769] The terminal accepts search queries from users via a GUI. The user enters a query (e.g., "latest trends in autonomous driving technology") into the text input field and clicks the search button.

[0770] Output: User's search query

[0771] Step 5: Query analysis and extraction of corresponding data

[0772] Input: User's search query

[0773] The server parses the search query received from the terminal and extracts keywords from the query using the Python spaCy library. Next, it searches the database using Elasticsearch to extract relevant technical documents and patent information. For example, it extracts the keyword "autonomous driving technology" from the query "latest trends in autonomous driving technology" and performs a search.

[0774] Output: Extracted technical literature and patent information

[0775] Step 6: Summary Generation

[0776] Input: Extracted technical documents and patent information

[0777] The server inputs the extracted data into a generative artificial intelligence model to generate a summary that includes key technical points, discoveries, and application examples. For example, it uses a GPT-3 model to quickly summarize information and provide it to the user.

[0778] Output: Summarized information

[0779] Step 7: Emotion Analysis

[0780] Input: User query input and operation information

[0781] The device collects emotional data from the user's text input and actions. For example, it uses an emotional analysis library (such as TextBlob or VADER) to determine if the user is experiencing emotions like "anxiety" or "interest." If the user types "difficult," the device analyzes that emotion.

[0782] Output: User's emotional state

[0783] Step 8: Customizing the information provided

[0784] Input: Summarized information and user's emotional state

[0785] The server customizes the summary information generated based on the analyzed emotional state. If the user is feeling "anxious," reassuring phrases and detailed information are added to the summary. For example, a sentence like, "This technology has already been securely implemented by major manufacturers and has a proven track record," might be added to the summary.

[0786] Output: Customized summary information

[0787] Step 9: Provide a summary

[0788] Input: Customized summary information

[0789] The server sends customized summary information to the terminal. The terminal displays the received summary information in its user interface, providing the information to the user.

[0790] Output: Summary information displayed to the user

[0791] Through the steps outlined above, this system enables the effective collection, analysis, and summarization of technical literature and intellectual property information, as well as the provision of information based on user sentiment. This allows users to quickly obtain the necessary information and make appropriate decisions.

[0792] (Application Example 2)

[0793] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0794] In today's world, users are required to acquire advanced information quickly and efficiently. However, a vast amount of technical literature and intellectual property information exists, and extracting and summarizing highly relevant information from it requires considerable time and effort. Furthermore, the information provided may lack effectiveness for the recipient because it is not appropriately customized according to the user's emotions and circumstances. In addition, advertising also faces the challenge of not providing appropriate information that is tailored to the user's situation.

[0795] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting technical literature and intellectual property information from a database, means for preprocessing the collected data and removing irrelevant information from the text, means for training a generative artificial intelligence model using the preprocessed data, means for customizing a summary generated based on the user's emotional state, and means for providing the user with advertising information including the customized summary. As a result, the user can efficiently obtain optimized technical literature and intellectual property information, and the receptivity of the information is improved because the most appropriate advertising information is provided according to the user's emotions and situation.

[0796] A "database" is a collection of information, including technical documents and intellectual property information. The information is stored in a searchable format.

[0797] "Technical documents" are materials that describe research results and technical explanations related to science and technology. This includes academic papers and technical reports.

[0798] "Intellectual property information" refers to information concerning intellectual property rights such as patents, trademarks, and copyrights. This includes patent application documents and registration certificates.

[0799] "Preprocessing" is the process of removing irrelevant information and noise from collected data and converting the data into an analyzable format. This also includes the process of removing HTML tags and special characters from text.

[0800] A "generative artificial intelligence model" is a machine learning model that has the ability to generate new information based on large amounts of data. Examples include BERT and GPT.

[0801] A "search query" is a request for input from a user to obtain specific information. It consists of keywords or questions.

[0802] "Summary generation" is the process of shortening long texts or data and extracting the main points. This is done automatically using artificial intelligence models.

[0803] "Emotional state" refers to the psychological state a user is in when receiving information. Specifically, it includes "excitement," "anxiety," and "interest."

[0804] "Customization" refers to adjusting and modifying information to suit the specific needs and feelings of the user.

[0805] "Advertising information" refers to information used to promote products and services. It is provided in an optimized format based on the user's interests and emotional state.

[0806] This invention relates to a system that collects technical literature and intellectual property information from a database and provides information optimized for the user. This system preprocesses the collected data, generates summaries using a generative artificial intelligence model, further customizes the information based on the user's emotional state, and provides advertising information along with it.

[0807] Hardware and software usage

[0808] The system is implemented using the following hardware and software:

[0809] Server: Responsible for large-scale data collection, preprocessing, training of generative artificial intelligence models, and summary generation.

[0810] Terminal: Responsible for receiving search queries from users and collecting sentiment data.

[0811] Generative artificial intelligence models: For example, BERT or GPT are used.

[0812] Emotion analysis engine: Emotion analysis is performed using the Hugging Face Transformers library.

[0813] Database: Stores technical literature and intellectual property information, and provides data via API.

[0814] Data processing and data calculation

[0815] Data Acquisition and Preprocessing

[0816] The server collects technical literature and intellectual property information from a database via an API. The collected data is then stored in JavaScript Object Notation (JSON) format and converted to text format. The text undergoes preprocessing such as noise reduction, tokenization, and stop word removal.

[0817] Training of generative artificial intelligence models

[0818] Using preprocessed data, the server trains a generative artificial intelligence model. This training is performed to generate summaries related to a specific technical domain. For example, PyTorch is used to train the model.

[0819] Emotion analysis

[0820] The device collects emotional data in real time as the user enters search queries. This emotional data is input into an emotional analysis engine through text analysis. The emotional analysis engine detects what emotional state the user is in, such as "excitement," "interest," or "anxiety."

[0821] Customized information provision

[0822] The server customizes the generated summary information based on sentiment analysis results. For example, it adds reassuring or attention-grabbing phrases to the summary. It then sends the customized summary along with advertising information to the device.

[0823] Specific example

[0824] When a user enters the query "latest smartphone technology," the server collects relevant information from technical literature and intellectual property databases. The collected data is then preprocessed and summarized by a generative artificial intelligence model. After the summary is generated, an emotion analysis engine analyzes the user's emotional state, and the summary information is customized based on the results.

[0825] For example, if a user expresses excitement, the summary information will include phrases that emphasize the appeal and future potential of the new technology. Finally, the customized summary information and related advertisements are sent to the device and displayed to the user.

[0826] Example of a prompt:

[0827] "Could you tell me about the latest trends in smartphone technology?"

[0828] The above describes the specific forms for carrying out the invention.

[0829] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0830] Step 1: Data Collection

[0831] The server collects relevant information from technical literature databases and intellectual property information databases via APIs. The input is a query containing keywords related to the technical literature and intellectual property information to be collected. The server stores the data collected via the API in JSON format. The output is a dataset of the collected technical literature and intellectual property information.

[0832] Step 2: Data Preprocessing

[0833] The server preprocesses the data collected in Step 1. Specifically, it removes noise such as HTML tags and special characters from the text and performs tokenization. The collected dataset of technical documents and intellectual property information is used as input. The output is clean data that has been de-noised and tokenized.

[0834] Step 3: Training the Generative AI Model

[0835] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using preprocessed data. The input is a preprocessed dataset. Training is performed using machine learning libraries such as PyTorch. The output is a generative artificial intelligence model trained to generate summaries corresponding to a specific technology domain.

[0836] Step 4: Accepting User Queries

[0837] The terminal receives search queries from the user through its interface. The input is the search query entered by the user (e.g., "Please tell me about the latest trends in smartphone technology."). The output is the received search query.

[0838] Step 5: Query analysis and data extraction

[0839] The server parses the search query received from the terminal and extracts keywords. It uses the user's search query as input. Next, the server extracts relevant technical literature and intellectual property information from the database based on the query. The output is a set of relevant technical literature and intellectual property information.

[0840] Step 6: Summary Generation

[0841] The server inputs extracted technical literature and intellectual property information into a generative artificial intelligence model to generate a summary. The extracted dataset and the trained generative AI model are used as input. The output is a summary containing the main points.

[0842] Step 7: Emotion Analysis

[0843] The terminal collects sentiment data when the user enters a search query. The input consists of user interface interactions and input text. The server inputs the collected sentiment data into a sentiment analysis engine to analyze the user's emotional state (e.g., "excited," "anxious"). The output is the user's emotional state.

[0844] Step 8: Customizing the information provided

[0845] The server customizes the generated summary information based on the analyzed user's emotional state. It uses the generated summary and the user's emotional state as input. For example, if the user is "excited," it adds more appealing expressions to the summary. The output is the customized summary information.

[0846] Step 9: Provide a summary

[0847] The server sends customized summary information and associated advertising information to the device. The device uses the customized summary information and generated advertising information as input. The device displays this information to the user. The output is the customized summary information and advertising information provided to the user.

[0848] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0849] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0850] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0851] [Third Embodiment]

[0852] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0853] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0854] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0855] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0856] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0857] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0858] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0859] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0860] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0861] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0862] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0863] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0864] This invention relates to a system for collecting technical documents and intellectual property information, enabling users to efficiently obtain this information. The system is characterized by collecting information from a database, performing preprocessing, generating a summary using a generative artificial intelligence model, and providing it to the user.

[0865] System Programs and Processing

[0866] 1. Data Collection

[0867] The server connects to technical literature databases and intellectual property databases to collect relevant information.

[0868] The server queries the database using keywords such as "autonomous driving" and retrieves the corresponding data.

[0869] The server stores the retrieved data in a structured format such as JSON.

[0870] 2. Data preprocessing

[0871] The server converts the collected data into text format.

[0872] The server cleans up the data by removing noise from the text (e.g., HTML tags and special characters).

[0873] The server tokenizes the clean data and removes irrelevant stop words.

[0874] The server saves the pre-processed data to the database.

[0875] 3. Training the AI ​​model

[0876] The server uses the pre-processed dataset to train a generative artificial intelligence model.

[0877] The server uses generative artificial intelligence models (e.g., BERT or GPT) to learn important patterns in the data.

[0878] The server stores the trained model and deploys it for inference.

[0879] 4. Acceptance of user queries

[0880] The terminal receives search queries from the user through the input interface.

[0881] The user enters keywords such as "latest information on autonomous driving technology."

[0882] 5. Query analysis and extraction of corresponding data

[0883] The server receives the search query sent from the terminal and parses the keywords.

[0884] The server extracts relevant technical literature and intellectual property information from the database based on the query.

[0885] 6. Summary Generation

[0886] The server inputs the extracted data into a generative artificial intelligence model and generates a summary.

[0887] The server verifies the summary results to ensure that the main points are reflected.

[0888] 7. Providing a summary

[0889] The server converts the generated summary into JSON format and sends it to the terminal.

[0890] The device displays summary information to the user.

[0891] Specific example

[0892] Data Acquisition and Preprocessing

[0893] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to autonomous driving. The collected data is stored in JSON format, and after noise reduction and tokenization, it is stored in the database.

[0894] Model training and query analysis

[0895] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize important information related to autonomous driving technology. When a user enters "latest information on autonomous driving technology" as a query, the server analyzes this query and extracts relevant papers and patents.

[0896] Summary generation and delivery

[0897] The extracted information is summarized by a generative artificial intelligence model. For example, a summary might be generated stating, "The latest autonomous driving technologies include AI-powered real-time road surface detection and high-precision obstacle detection using laser sensors." The generated summary is sent from the server to the terminal, where the user can view it through the interface.

[0898] Based on the above, the present invention is a system that enables users to effectively collect and summarize technical documents and intellectual property information.

[0899] The following describes the processing flow.

[0900] Step 1:

[0901] The server connects to technical literature databases and intellectual property databases. For example, it uses APIs to collect data from the "technical literature database" and the "intellectual property database."

[0902] Step 2:

[0903] The server converts the collected data into JSON format and saves it to local or cloud storage. The collected data includes the title of the paper, authors, abstracts, patent numbers, and summaries of the inventions.

[0904] Step 3:

[0905] The server preprocesses the collected data. Specifically, it removes HTML tags and special characters from the text. Next, it tokenizes the cleansed data and removes unwanted words (stop words).

[0906] Step 4:

[0907] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model to begin the learning process.

[0908] Step 5:

[0909] The server evaluates the trained generative artificial intelligence model on a test dataset to verify its accuracy. Once the evaluation is complete, the model is deployed for inference.

[0910] Step 6:

[0911] The terminal receives search queries from users through its interface. For example, it might accept a query such as "latest trends in autonomous driving technology."

[0912] Step 7:

[0913] The server analyzes the search query received from the terminal and extracts keywords. For example, "autonomous driving technology" and "latest trends" might be extracted as keywords.

[0914] Step 8:

[0915] The server extracts relevant technical literature and intellectual property information from the database based on the analyzed query. The extracted data is temporarily loaded into memory.

[0916] Step 9:

[0917] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The summary includes key technical points, findings, and application examples.

[0918] Step 10:

[0919] The server converts the generated summary into JSON format and sends it to the terminal. Because the transmitted summary data is structured, it is displayed in a user-friendly format.

[0920] Step 11:

[0921] The terminal displays the received summary results in the user interface. Users can review the summary information corresponding to their search queries and quickly obtain the necessary technical information.

[0922] This series of processes allows users to efficiently collect the latest technical literature and intellectual property information, and obtain the necessary information in a summarized format.

[0923] (Example 1)

[0924] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0925] Because technical literature and intellectual property information are vast, it is difficult for individual users to efficiently collect the necessary information and summarize and understand the key points. Traditional systems primarily rely on manual information gathering and summarization, which is time-consuming and labor-intensive. Furthermore, the process of extracting and summarizing appropriate information based on search queries is often inefficient. Therefore, there is a need for a system that can automatically collect, preprocess, and summarize technical literature and intellectual property information, and efficiently provide it to users.

[0926] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0927] In this invention, the server includes means for collecting relevant information from a technical literature database and an intellectual property database; means for converting the collected information into text format, removing noise, and creating clean data; and means for tokenizing the clean data and preprocessing irrelevant information. This makes it possible to automatically collect and preprocess technical literature and intellectual property information to generate a dataset.

[0928] The server also includes means for training a generative artificial intelligence model using a preprocessed dataset, means for receiving search queries from users, means for analyzing the received search queries and extracting relevant information from technical literature databases and intellectual property databases, means for generating summaries of the extracted information using the generative artificial intelligence model, and means for converting the generated summaries into JSON format and providing them to users. This makes it possible to efficiently extract relevant information, summarize the key points, and provide them to users.

[0929] A "technical literature database" is a database that collects technical information, such as academic papers and technical reports, and stores it in a searchable format.

[0930] An "intellectual property database" is a database that collects and stores information related to intellectual property, such as patents, trademarks, and copyrights, in a searchable format.

[0931] A "generative artificial intelligence model" is an artificial intelligence model that has the ability to generate new data in tasks such as natural language processing and image generation.

[0932] "JSON format" is an abbreviation for JavaScript Object Notation, and it is a lightweight data exchange format for structuring and storing data.

[0933] "Noise reduction" is the process of removing unnecessary information (e.g., HTML tags and special characters) from data to generate clean data.

[0934] A "query" is a type of inquiry that a user enters in a database or search system to retrieve specific information.

[0935] "Tokenization" is the process of dividing text data into semantic units (tokens) such as words and phrases.

[0936] "Preprocessing" refers to the process of preparing data, such as cleaning up and standardizing its format, before analyzing the data or training models.

[0937] A "summary" is information that extracts the important points from a vast amount of information and presents them in a short format.

[0938] "Analysis" is the process of analyzing data and query content to understand and process their meaning and structure.

[0939] "Extraction" is the process of retrieving data or information based on specific conditions.

[0940] This invention relates to a system for efficiently collecting and processing technical literature and intellectual property information and providing it to users in a summarized format. The system collects relevant information from technical literature databases and intellectual property databases, performs preprocessing, generates summaries using a generative artificial intelligence model, and provides them to users.

[0941] Data collection

[0942] The server periodically accesses technical literature databases (e.g., IEEE Xplore) and intellectual property databases (e.g., USPTO) to collect relevant information. For example, the server automatically sends queries at 3:00 AM every day to retrieve the latest papers and patent information related to keywords such as "autonomous driving" and "AI technology." The collected information is temporarily stored in data storage in a structured JSON format.

[0943] Data preprocessing

[0944] The server converts the collected data into text format and removes noise. First, HTML tags and special characters are removed using regular expressions. Next, the text data is tokenized (divided into words and phrases), and irrelevant stop words (e.g., "a" and "the" in English) are removed using the NLTK library. The pre-processed data is then saved back to the database.

[0945] Training an AI model

[0946] The server trains a generative artificial intelligence (AI) model using a pre-processed dataset. For example, it trains a BERT model using the Hugging Face Transformers library. The server uses 80% of the data for training and 20% for validation. Once the training is complete, the model is stored in the Hugging Face model hub.

[0947] Acceptance of user queries

[0948] The terminal receives search queries from the user through an input interface. For example, it might be implemented as a web application running in a browser, where the user enters a query such as "latest autonomous driving technology" into the search bar. This query is sent to the server in real time.

[0949] Query analysis and extraction of corresponding data

[0950] The server receives search queries sent from the terminal and analyzes the keywords using TF-IDF (Term Frequency-Inverse Document Frequency). Based on the analysis results, the server extracts relevant information from the database. For example, it might pick out literature and patents related to "latest autonomous driving technology."

[0951] Summary generation

[0952] The server inputs the extracted literature and patent information into a generative artificial intelligence model to generate a summary. Specifically, it feeds the extracted text into a BERT model to generate a summarized text. During this process, the model undergoes multiple validations to ensure that all important points are included.

[0953] Summary

[0954] The server converts the generated summary into JSON format and sends it to the terminal. The terminal displays this summary information to the user. For example, in the user interface of a web application, the summary text is displayed in an easy-to-read format. The user can review the displayed summary and access detailed information as needed.

[0955] Specific example

[0956] Data Acquisition and Preprocessing

[0957] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to "autonomous driving." The collected data is stored in JSON format and, after preprocessing such as removing HTML tags and special characters and tokenizing the text, is stored in the database.

[0958] Model training and query analysis

[0959] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize key information related to "autonomous driving technology." When a user enters "latest information on autonomous driving technology" as a query, the server analyzes this query and extracts relevant papers and patents.

[0960] Summary generation and delivery

[0961] The extracted information is summarized by a generative artificial intelligence model. For example, a summary might be generated stating, "The latest autonomous driving technologies include AI-powered real-time road surface detection and high-precision obstacle detection using laser sensors." The generated summary is sent from the server to the terminal, where the user can view it in a web browser.

[0962] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0963] Step 1: Data collection from the database

[0964] The server connects to technical literature databases and intellectual property databases to collect information related to specific keywords (e.g., "autonomous driving" or "AI technology"). The input is a keyword query, and the output is the collected data in JSON format. Specifically, the server automatically sends queries at 3 AM every day and temporarily stores the retrieved papers and patent information in data storage.

[0965] Step 2: Data Preprocessing

[0966] The server converts the collected data into text format and removes noise. The input is collected data in JSON format, and the output is clean text data. Specifically, the server uses regular expressions to remove HTML tags and special characters, then tokenizes the text data into words and phrases, and removes irrelevant stop words using the NLTK library. This clean data is then stored again in the database.

[0967] Step 3: Training the AI ​​model

[0968] The server trains a generative artificial intelligence model using a preprocessed dataset. The input is a preprocessed dataset, and the output is a trained generative AI model. Specifically, the server trains a BERT model using the Hugging Face Transformers library, using 80% of the data for training and 20% for validation. The trained model is stored in the Hugging Face model hub.

[0969] Step 4: Accepting User Queries

[0970] The terminal receives search queries from the user through an input interface. The input is the user's search query (e.g., "latest autonomous driving technology"), and the output is that query being sent to the server. Specifically, the user enters the query into the search bar of the web application, and that query is sent to the server in real time.

[0971] Step 5: Query analysis and extraction of corresponding data

[0972] The server receives search queries sent from terminals and parses the keywords. The input is the user's search query, and the output is relevant information extracted from the database. Specifically, the server uses TF-IDF (Term Frequency-Inverse Document Frequency) to parse the query and extracts relevant technical documents and patent information from the database based on the results.

[0973] Step 6: Summary Generation

[0974] The server inputs the extracted information into a generative artificial intelligence model to generate a summary. The input includes extracted technical documents and patent information, and the output is the generated summary. Specifically, the server feeds the extracted text data into a BERT model to generate the summary. During this process, the model undergoes multiple validations to ensure that the key points are reflected.

[0975] Step 7: Provide a summary

[0976] The server converts the generated summary into JSON format and sends it to the terminal. The input is the generated summary text, and the output is the summary data sent to the terminal in JSON format. Specifically, the server converts the summary into JSON format and provides it to the user through the web application's user interface. The terminal displays the received summary to the user, who can access detailed information as needed.

[0977] (Application Example 1)

[0978] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0979] Traditionally, the collection and summarization of technical literature and intellectual property information was often done manually, requiring considerable time and effort. Furthermore, there was no way for users to instantly access the latest technical information about products within a virtual store, making effective information gathering and provision difficult. This resulted in reduced user convenience and impaired the efficiency of information gathering.

[0980] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0981] In this invention, the server includes means for collecting technical literature and intellectual property information from a database; means for preprocessing the collected data and removing irrelevant information from the text; means for training a generative artificial intelligence model using the preprocessed data; means for receiving search queries from users; means for extracting relevant information from the database based on the received search queries; means for generating a summary of the extracted information using the generative artificial intelligence model; means for providing the generated summary to the user; and means for enabling the user to instantly obtain the latest technical and intellectual property information about products in a virtual store using a smartphone or smart glasses. This makes it possible for users to efficiently collect technical literature and intellectual property information and instantly check the summarized information.

[0982] A "database" is a system for organizing, systematizing, and centrally managing information.

[0983] "Technical literature" refers to documents that describe research results and knowledge related to science and technology.

[0984] "Intellectual property information" refers to information concerning intellectual property rights such as patents, trademarks, and copyrights.

[0985] "Preprocessing" is the process of removing noise and irrelevant information from raw data and converting it into an analyzable format.

[0986] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates text and other data based on training data.

[0987] A "search query" is a keyword or phrase that a user uses to search for specific information.

[0988] A "summary" is a short, concise version of a longer text, extracting the most important information from it.

[0989] A "smartphone" is a portable device that, in addition to telephone functionality, also possesses advanced computer capabilities.

[0990] "Smart glasses" are glasses-type devices that display information in the field of view and have the ability to connect to the internet.

[0991] A "virtual store" is an online shop that exists on the internet, where users can browse and purchase products in a virtual space.

[0992] A "product" is an item or service that is manufactured and supplied to consumers.

[0993] This invention relates to a system that collects technical literature and intellectual property information and provides summarized information using a generative artificial intelligence model. This system enables users, particularly in virtual stores, to instantly obtain the latest technical and intellectual property information about products using a smartphone or smart glasses.

[0994] System Overview

[0995] 1. Collection from databases

[0996] The server connects to technical literature databases and intellectual property databases to collect relevant information. For example, it queries the databases using keywords such as "autonomous driving" and retrieves corresponding data. The retrieved data is stored in a structured format such as JSON.

[0997] 2. Data preprocessing

[0998] The server converts the collected data into text format. It then cleans the data by removing noise (e.g., HTML tags and special characters). Next, it tokenizes the text data and removes irrelevant stop words. The pre-processed data is then stored back in the database.

[0999] 3. Training the AI ​​model

[1000] The server uses preprocessed data to train a generative artificial intelligence model. This model may be one such model, such as BERT or GPT. Once trained, the model learns important patterns and is deployed for inference.

[1001] 4. Acceptance of user queries

[1002] The terminal receives search queries from users through an input interface. For example, a user might enter keywords such as "latest information on autonomous driving technology."

[1003] 5. Query analysis and extraction of corresponding data

[1004] The server receives search queries sent from the terminal and analyzes the keywords. Next, it extracts relevant technical literature and intellectual property information from the database based on the query.

[1005] 6. Summary Generation

[1006] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The generated summary is then verified to ensure that the main points are reflected before being provided to the user.

[1007] 7. Providing a summary

[1008] The server sends the summarized results, converted to JSON format, to the terminal. The terminal then displays the generated summarized information to the user.

[1009] Specific example

[1010] For example, if the user wants to obtain information about "autonomous driving," the server sends queries to technical literature databases and intellectual property databases, collects papers and patents related to autonomous driving, and generates clean data. A trained generative artificial intelligence model generates a summary from this data, such as "the latest autonomous driving technologies include AI-based real-time road surface detection and high-precision obstacle detection using laser sensors." Users can instantly check this summary information via their smartphone or smart glasses. This invention enables rapid information provision within virtual stores.

[1011] Example of a prompt

[1012] The following prompts are used as input to the generative artificial intelligence model:

[1013] Summarize the latest information on autonomous driving technology.

[1014] The above describes the embodiments of the present invention. This system allows users to efficiently collect and verify the latest technical and intellectual property information within a virtual store.

[1015] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1016] Step 1:

[1017] The server connects to technical literature databases and intellectual property databases to collect relevant information. The server queries the databases using keywords (e.g., "autonomous driving") and stores the retrieved data in JSON format.

[1018] Input: Keyword query

[1019] Output: Technical document data and intellectual property information in JSON format

[1020] Specific operation: The server accesses the database via an API and retrieves and saves the relevant data.

[1021] Step 2:

[1022] The server preprocesses the collected data. Specifically, it converts it to text data, removes noise (e.g., HTML tags and special characters), and removes irrelevant stop words. It also performs tokenization. The preprocessed data is then stored again in the database.

[1023] Input: Technical document data and intellectual property information in JSON format

[1024] Output: Preprocessed text data

[1025] Specific operation: The server performs text data cleanup, tokenization, and stop word removal.

[1026] Step 3:

[1027] The server trains generative artificial intelligence models using preprocessed data. Specifically, it inputs data into models such as BERT and GPT to learn patterns. Once the training is complete, the models are stored on the server and deployed for inference.

[1028] Input: Preprocessed text data

[1029] Output: Trained generative artificial intelligence model

[1030] Specific operation: The server inputs data into the model and executes and completes the training process.

[1031] Step 4:

[1032] The terminal receives search queries from the user. The user enters search keywords (e.g., "latest information on autonomous driving technology") through the interface.

[1033] Input: Search keyword (user query)

[1034] Output: Search query (sent from terminal to server)

[1035] Specific operation: The terminal receives queries through the user interface and sends them to the server.

[1036] Step 5:

[1037] The server receives search queries sent from terminals and analyzes the keywords. Based on the queries, the server extracts relevant technical literature and intellectual property information from its database.

[1038] Input: Search query from the user

[1039] Output: Extracted technical literature and intellectual property information

[1040] Specific operation: The server parses the query and quickly extracts relevant information from the database.

[1041] Step 6:

[1042] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The generated summary is then verified to ensure that the main points are reflected before being provided to the user.

[1043] Input: Extracted technical literature and intellectual property information

[1044] Output: Summarized information

[1045] Specific operation: The server uses a trained generative artificial intelligence model to generate and verify a summary of the extracted data.

[1046] Step 7:

[1047] The server converts the generated summary into JSON format and sends it to the terminal. The terminal then displays the summary information to the user.

[1048] Input: Summary information (generated by a generative AI model)

[1049] Output: Summary information in JSON format (sent from server to terminal)

[1050] Specific operation: The server encodes the summary information and sends it to the terminal. The terminal displays the summary information to the user.

[1051] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1052] This invention relates to a system for collecting technical literature and intellectual property information, enabling users to efficiently obtain this information. The system is characterized by collecting information from a database, performing preprocessing, generating summaries using a generative artificial intelligence model, and providing them to the user. Furthermore, it incorporates an emotion engine that analyzes the user's emotions, and adjusts and customizes the information based on the user's emotional state.

[1053] System Programs and Processing

[1054] 1. Data Collection

[1055] The server connects to technical literature databases and intellectual property databases to collect relevant information. For example, it uses APIs to retrieve "technical literature" and "intellectual property information" from the databases.

[1056] 2. Data preprocessing

[1057] The server converts the collected data into text format and removes noise (e.g., HTML tags and special characters). Next, it tokenizes the clean data and removes irrelevant stop words.

[1058] 3. Training the AI ​​model

[1059] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model, and the learning process proceeds.

[1060] 4. Acceptance of user queries

[1061] The terminal receives search queries from the user through its interface. For example, the user might enter the query, "Latest trends in autonomous driving technology."

[1062] 5. Query analysis and extraction of corresponding data

[1063] The server analyzes the search query received from the terminal and extracts keywords. Next, it extracts relevant technical literature and intellectual property information from the database based on the query.

[1064] 6. Summary Generation

[1065] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The summary includes key technical points, findings, and application examples.

[1066] 7. Emotion analysis

[1067] The device collects emotional data from the user's interface operations and input.

[1068] The server uses an emotion engine to analyze the user's emotions. For example, it uses text analysis to determine whether the user is experiencing emotions such as "anxiety" or "interest."

[1069] 8. Customization of information provision

[1070] The server customizes the generated summary information based on the analyzed emotional state. For example, if the user is feeling "anxious," it adds reassuring phrases.

[1071] 9. Providing a summary

[1072] The server sends a customized summary to the terminal.

[1073] The device displays summary information to the user.

[1074] Specific example

[1075] Data Acquisition and Preprocessing

[1076] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to autonomous driving. The collected data is stored in JSON format, and after noise reduction and tokenization, it is stored in the database.

[1077] Model training and query analysis

[1078] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize important information related to autonomous driving technology. When a user enters "latest information on autonomous driving technology" as a query, the server analyzes this query and extracts relevant papers and patents.

[1079] Summarization generation and sentiment analysis

[1080] The extracted information is summarized by a generative artificial intelligence model. Simultaneously, the terminal collects the user's emotions when they enter queries, and the server analyzes those emotions using an emotion engine.

[1081] Customization and delivery

[1082] The summarized information is customized based on the user's emotional state. For example, if the user indicates "anxiety," the summary will include expressions that emphasize stability and reliability. Finally, the customized summary information is sent from the server to the terminal and displayed to the user.

[1083] In summary, the present invention is a system that enables users to effectively collect and summarize technical literature and intellectual property information, as well as to provide customized information based on the user's emotional state.

[1084] The following describes the processing flow.

[1085] Step 1:

[1086] The server connects to technical literature databases and intellectual property databases, and uses APIs and queries to collect "technical literature" and "intellectual property information." For example, it retrieves data related to "autonomous driving."

[1087] Step 2:

[1088] The server converts the collected data into JSON format and stores it securely in local or cloud storage.

[1089] Step 3:

[1090] The server preprocesses the collected data. Specifically, it removes HTML tags and special characters from the text. Furthermore, it performs natural language processing such as tokenization and stop word removal to clean the data.

[1091] Step 4:

[1092] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model to learn important patterns and features.

[1093] Step 5:

[1094] The server evaluates the trained generative artificial intelligence model on a test dataset to verify its accuracy. Once the evaluation is complete, the model is deployed for production use.

[1095] Step 6:

[1096] The terminal receives search queries from the user through its interface. For example, the user enters the query "latest trends in autonomous driving technology."

[1097] Step 7:

[1098] The terminal collects emotional data from the user's facial expressions, typing speed, and context when they enter queries. The collected emotional data is sent to the server.

[1099] Step 8:

[1100] The server analyzes the search queries received from the terminal and extracts "autonomous driving technology" and "latest trends" as keywords. Next, it extracts technical literature and intellectual property information related to these keywords from the database.

[1101] Step 9:

[1102] The server inputs the extracted information into a generative artificial intelligence model to generate a summary. The summary includes key technical points, findings, and application examples.

[1103] Step 10:

[1104] The server uses an emotion engine to analyze the transmitted emotion data. For example, it might identify the user's emotion as "anxiety."

[1105] Step 11:

[1106] The server customizes the summary generated based on the sentiment analysis results. For example, if "anxiety" is identified, it adds reassuring wording to the summary.

[1107] Step 12:

[1108] The server converts the customized summary into JSON format and sends it to the terminal.

[1109] Step 13:

[1110] The device displays the received, customized summary results in the user interface. Users can review the summary information for their queries and obtain information tailored to their emotional state.

[1111] This series of processes allows users to efficiently collect the latest technical literature and intellectual property information and confidently utilize the summarized information. The system also takes the user's emotional state into account, enabling more personalized information delivery.

[1112] (Example 2)

[1113] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1114] In conventional technical information gathering systems, it was difficult to efficiently collect relevant information from vast databases and summarize that information. Furthermore, the summarized information did not adapt to the user's emotional state, resulting in a decline in the quality of information provided. As a result, users were unable to obtain the necessary information quickly and appropriately, leading to wasted effort and time.

[1115] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1116] In this invention, the server includes means for collecting technical documents and intellectual property information from a database, means for preprocessing the collected data and removing irrelevant information from the text, and means for training a generative artificial intelligence model using the preprocessed data. This makes it possible to efficiently preprocess the collected data and train the generative artificial intelligence model with high accuracy. Furthermore, by including means for extracting relevant information from the database based on a received search query and generating a summary of that information using the generative artificial intelligence model, means for analyzing the user's emotional state during summary generation and customizing the summary information based on that state, and means for providing the generated summary to the user, it becomes possible to provide high-quality information adapted to the user's emotions, thereby improving the efficiency and accuracy of information collection.

[1117] A "database" is an information system that systematically stores multiple pieces of data, making it possible to search for and retrieve them.

[1118] "Technical literature" refers to documents such as academic papers, research reports, and technical reports related to a specific technical field.

[1119] "Intellectual property information" refers to data and documents related to intellectual property rights, such as patents, trademarks, and copyrights.

[1120] "Preprocessing" refers to a series of operations that transform data into a format suitable for analysis and model training.

[1121] A "generative artificial intelligence model" is an artificial intelligence model that can generate new text or information based on input data.

[1122] "Tokenization" is the process of dividing text data into meaningful units such as words or phrases.

[1123] A "stop word" is a common word that is considered low in importance and is excluded in text analysis and natural language processing.

[1124] A "search query" is a keyword or phrase that a user enters to search for information.

[1125] A "summary" is a shortened version of a document or piece of information, containing the main points extracted from it.

[1126] "Sentiment analysis" is the process of determining a user's emotional state from their text and behavior.

[1127] "Customization" refers to modifying or adjusting information and services to suit the specific needs and circumstances of the user.

[1128] "User interface" is a general term for the screens and operating methods that users use to interact with a system.

[1129] "Noise" refers to data or information that is unnecessary or harmful to the analysis.

[1130] Hugging Face is an organization that provides open-source libraries and tools for natural language processing.

[1131] The "Transformers library" is an open-source software library used for building and training generative artificial intelligence models.

[1132] Elasticsearch is a full-text search engine and analytics engine, a system designed for high-speed searching of large amounts of data.

[1133] "spaCy" is a high-performance natural language processing (NLP) library used for text analysis and model building.

[1134] "BeautifulSoup" is a Python library for parsing HTML and XML files and extracting specific information.

[1135] TextBlob is a Python library for easily performing sentiment analysis and classification of text data.

[1136] "VADER" is a Python library specifically designed for text sentiment analysis.

[1137] This invention relates to a system for efficiently collecting, summarizing, and providing technical literature and intellectual property information to users. This system is characterized by collecting information from a database, performing preprocessing, generating summaries using a generative artificial intelligence model, and providing information customized based on the user's sentiment.

[1138] Data collection

[1139] The server connects to technical literature databases and intellectual property databases via APIs to collect relevant information. For example, it collects papers and patent data related to "autonomous driving" from databases such as IEEE Xplore and Google Patents. The server authenticates using an API key, sends search queries, and retrieves data.

[1140] Specific example: The server sends "self-driving cars" as a search query to the IEEE Xplore API to retrieve a list of papers related to autonomous driving.

[1141] Data preprocessing

[1142] The server converts the collected data into text format and removes noise such as HTML tags and special characters using Python's BeautifulSoup library. Next, the clean text data is tokenized, and irrelevant stop words are removed using Python's NLTK library.

[1143] Specific example: Perform the operation of removing HTML tags from acquired research paper data and extracting only the text portion. In the text "The advancements in self-driving cars include sensors and AI technology.", stop words such as "The" and "in" are removed, and the tokens "advancements, self-driving, cars, include, sensors, AI, technology" are extracted.

[1144] Training an AI model

[1145] The server uses preprocessed data to train generative artificial intelligence models (e.g., BERT or GPT). The Hugging Face Transformers library is used to train the generative AI models. The GPT-3 model is trained using preprocessed text data on autonomous driving technology, aiming for highly accurate summary generation.

[1146] Acceptance of user queries

[1147] The terminal accepts search queries from users through a GUI (Graphical User Interface). For example, a user might enter the query "latest trends in autonomous driving technology" and click the search button.

[1148] Specific example: The user enters "latest trends in autonomous driving technology" and presses the search button.

[1149] Query analysis and extraction of corresponding data

[1150] The server parses the search query received from the terminal and extracts keywords from the query using the Python spaCy library. Next, it searches the database using Elasticsearch and extracts relevant technical documents and patent information.

[1151] Specific example: Extract the keyword "autonomous driving technology" from the query "latest trends in autonomous driving technology," and use Elasticsearch to search for papers and patent information related to this keyword.

[1152] Summary generation

[1153] The server inputs the extracted data into a generative artificial intelligence model to generate a summary that includes key technical points, discoveries, and application examples. For example, it can use a GPT-3 model to quickly summarize information and provide it to the user.

[1154] Specific example: Input extracted technical documents and patent information into GPT-3 to generate a "Summary of the Latest Trends in Autonomous Driving Technology."

[1155] Emotion analysis

[1156] The device collects emotional data from the user's text input and actions. For example, it uses an emotional analysis library (e.g., TextBlob or VADER) to analyze whether the user is experiencing emotions such as "anxiety" or "interest."

[1157] Specific example: If a user enters a keyword like "difficult," the TextBlob is used to analyze the emotion "anxiety."

[1158] Customization of information provision

[1159] The server customizes the summary information generated based on the analyzed emotional state. For example, if the user is feeling "anxious," it adds reassuring phrases and detailed information.

[1160] Specific example: Add a sentence to the summary such as, "This technology has already been safely implemented by major manufacturers and has a proven track record."

[1161] Summary

[1162] The server sends customized summary information to the terminal, which then displays it to the user.

[1163] Specific example: Display a customized summary in the user interface of a browser.

[1164] Through the steps described above, this system can efficiently collect and analyze technical literature and intellectual property information, and provide information tailored to the user's emotions. This allows users to quickly obtain the necessary information and deepen their understanding.

[1165] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1166] Program processing flow

[1167] Step 1: Data Collection

[1168] Input: Search queries to technical literature databases and intellectual property databases

[1169] The server connects to technical literature databases and intellectual property databases via APIs and collects data related to the information requested by the user. In this process, the server authenticates using an API key and sends search queries to retrieve data. For example, it might send "self-driving cars" as a query to the IEEE Xplore API.

[1170] Output: Collected technical literature data and intellectual property data

[1171] Step 2: Data Preprocessing

[1172] Input: Collected technical literature data and intellectual property data

[1173] The server converts the collected data into text format and removes noise such as HTML tags and special characters. The Python BeautifulSoup library is used for this process. Furthermore, the clean text data is tokenized, and irrelevant stop words are removed using the NLTK library. In this way, clean data suitable for analysis is obtained.

[1174] Output: Pre-processed clean data

[1175] Step 3: Training the AI ​​model

[1176] Input: Pre-processed clean data

[1177] The server uses preprocessed data to train generative artificial intelligence models (e.g., BERT or GPT). This training utilizes the Hugging Face Transformers library. The generative AI models learn to generate highly accurate summaries based on a large amount of technical literature data.

[1178] Output: Trained generative AI model

[1179] Step 4: Accepting User Queries

[1180] Input: User's search query

[1181] The terminal accepts search queries from users via a GUI. The user enters a query (e.g., "latest trends in autonomous driving technology") into the text input field and clicks the search button.

[1182] Output: User's search query

[1183] Step 5: Query analysis and extraction of corresponding data

[1184] Input: User's search query

[1185] The server parses the search query received from the terminal and extracts keywords from the query using the Python spaCy library. Next, it searches the database using Elasticsearch to extract relevant technical documents and patent information. For example, it extracts the keyword "autonomous driving technology" from the query "latest trends in autonomous driving technology" and performs a search.

[1186] Output: Extracted technical literature and patent information

[1187] Step 6: Summary Generation

[1188] Input: Extracted technical documents and patent information

[1189] The server inputs the extracted data into a generative artificial intelligence model to generate a summary that includes key technical points, discoveries, and application examples. For example, it uses a GPT-3 model to quickly summarize information and provide it to the user.

[1190] Output: Summarized information

[1191] Step 7: Emotion Analysis

[1192] Input: User query input and operation information

[1193] The device collects emotional data from the user's text input and actions. For example, it uses an emotional analysis library (such as TextBlob or VADER) to determine if the user is experiencing emotions like "anxiety" or "interest." If the user types "difficult," the device analyzes that emotion.

[1194] Output: User's emotional state

[1195] Step 8: Customizing the information provided

[1196] Input: Summarized information and user's emotional state

[1197] The server customizes the summary information generated based on the analyzed emotional state. If the user is feeling "anxious," reassuring phrases and detailed information are added to the summary. For example, a sentence like, "This technology has already been securely implemented by major manufacturers and has a proven track record," might be added to the summary.

[1198] Output: Customized summary information

[1199] Step 9: Provide a summary

[1200] Input: Customized summary information

[1201] The server sends customized summary information to the terminal. The terminal displays the received summary information in its user interface, providing the information to the user.

[1202] Output: Summary information displayed to the user

[1203] Through the steps outlined above, this system enables the effective collection, analysis, and summarization of technical literature and intellectual property information, as well as the provision of information based on user sentiment. This allows users to quickly obtain the necessary information and make appropriate decisions.

[1204] (Application Example 2)

[1205] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1206] In today's world, users are required to acquire advanced information quickly and efficiently. However, a vast amount of technical literature and intellectual property information exists, and extracting and summarizing highly relevant information from it requires considerable time and effort. Furthermore, the information provided may lack effectiveness for the recipient because it is not appropriately customized according to the user's emotions and circumstances. In addition, advertising also faces the challenge of not providing appropriate information that is tailored to the user's situation.

[1207] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting technical literature and intellectual property information from a database, means for preprocessing the collected data and removing irrelevant information from the text, means for training a generative artificial intelligence model using the preprocessed data, means for customizing a summary generated based on the user's emotional state, and means for providing the user with advertising information including the customized summary. As a result, the user can efficiently obtain optimized technical literature and intellectual property information, and the receptivity of the information is improved because the most appropriate advertising information is provided according to the user's emotions and situation.

[1208] A "database" is a collection of information, including technical documents and intellectual property information. The information is stored in a searchable format.

[1209] "Technical documents" are materials that describe research results and technical explanations related to science and technology. This includes academic papers and technical reports.

[1210] "Intellectual property information" refers to information concerning intellectual property rights such as patents, trademarks, and copyrights. This includes patent application documents and registration certificates.

[1211] "Preprocessing" is the process of removing irrelevant information and noise from collected data and converting the data into an analyzable format. This also includes the process of removing HTML tags and special characters from text.

[1212] A "generative artificial intelligence model" is a machine learning model that has the ability to generate new information based on large amounts of data. Examples include BERT and GPT.

[1213] A "search query" is a request for input from a user to obtain specific information. It consists of keywords or questions.

[1214] "Summary generation" is the process of shortening long texts or data and extracting the main points. This is done automatically using artificial intelligence models.

[1215] "Emotional state" refers to the psychological state a user is in when receiving information. Specifically, it includes "excitement," "anxiety," and "interest."

[1216] "Customization" refers to adjusting and modifying information to suit the specific needs and feelings of the user.

[1217] "Advertising information" refers to information used to promote products and services. It is provided in an optimized format based on the user's interests and emotional state.

[1218] This invention relates to a system that collects technical literature and intellectual property information from a database and provides information optimized for the user. This system preprocesses the collected data, generates summaries using a generative artificial intelligence model, further customizes the information based on the user's emotional state, and provides advertising information along with it.

[1219] Hardware and software usage

[1220] The system is implemented using the following hardware and software:

[1221] Server: Responsible for large-scale data collection, preprocessing, training of generative artificial intelligence models, and summary generation.

[1222] Terminal: Responsible for receiving search queries from users and collecting sentiment data.

[1223] Generative artificial intelligence models: For example, BERT or GPT are used.

[1224] Emotion analysis engine: Emotion analysis is performed using the Hugging Face Transformers library.

[1225] Database: Stores technical literature and intellectual property information, and provides data via API.

[1226] Data processing and data calculation

[1227] Data Acquisition and Preprocessing

[1228] The server collects technical literature and intellectual property information from a database via an API. The collected data is then stored in JavaScript Object Notation (JSON) format and converted to text format. The text undergoes preprocessing such as noise reduction, tokenization, and stop word removal.

[1229] Training of generative artificial intelligence models

[1230] Using preprocessed data, the server trains a generative artificial intelligence model. This training is performed to generate summaries related to a specific technical domain. For example, PyTorch is used to train the model.

[1231] Emotion analysis

[1232] The device collects emotional data in real time as the user enters search queries. This emotional data is input into an emotional analysis engine through text analysis. The emotional analysis engine detects what emotional state the user is in, such as "excitement," "interest," or "anxiety."

[1233] Customized information provision

[1234] The server customizes the generated summary information based on sentiment analysis results. For example, it adds reassuring or attention-grabbing phrases to the summary. It then sends the customized summary along with advertising information to the device.

[1235] Specific example

[1236] When a user enters the query "latest smartphone technology," the server collects relevant information from technical literature and intellectual property databases. The collected data is then preprocessed and summarized by a generative artificial intelligence model. After the summary is generated, an emotion analysis engine analyzes the user's emotional state, and the summary information is customized based on the results.

[1237] For example, if a user expresses excitement, the summary information will include phrases that emphasize the appeal and future potential of the new technology. Finally, the customized summary information and related advertisements are sent to the device and displayed to the user.

[1238] Example of a prompt:

[1239] "Could you tell me about the latest trends in smartphone technology?"

[1240] The above describes the specific forms for carrying out the invention.

[1241] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1242] Step 1: Data Collection

[1243] The server collects relevant information from technical literature databases and intellectual property information databases via APIs. The input is a query containing keywords related to the technical literature and intellectual property information to be collected. The server stores the data collected via the API in JSON format. The output is a dataset of the collected technical literature and intellectual property information.

[1244] Step 2: Data Preprocessing

[1245] The server preprocesses the data collected in Step 1. Specifically, it removes noise such as HTML tags and special characters from the text and performs tokenization. The collected dataset of technical documents and intellectual property information is used as input. The output is clean data that has been de-noised and tokenized.

[1246] Step 3: Training the Generative AI Model

[1247] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using preprocessed data. The input is a preprocessed dataset. Training is performed using machine learning libraries such as PyTorch. The output is a generative artificial intelligence model trained to generate summaries corresponding to a specific technology domain.

[1248] Step 4: Accepting User Queries

[1249] The terminal receives search queries from the user through its interface. The input is the search query entered by the user (e.g., "Please tell me about the latest trends in smartphone technology."). The output is the received search query.

[1250] Step 5: Query analysis and data extraction

[1251] The server parses the search query received from the terminal and extracts keywords. It uses the user's search query as input. Next, the server extracts relevant technical literature and intellectual property information from the database based on the query. The output is a set of relevant technical literature and intellectual property information.

[1252] Step 6: Summary Generation

[1253] The server inputs extracted technical literature and intellectual property information into a generative artificial intelligence model to generate a summary. The extracted dataset and the trained generative AI model are used as input. The output is a summary containing the main points.

[1254] Step 7: Emotion Analysis

[1255] The terminal collects sentiment data when the user enters a search query. The input consists of user interface interactions and input text. The server inputs the collected sentiment data into a sentiment analysis engine to analyze the user's emotional state (e.g., "excited," "anxious"). The output is the user's emotional state.

[1256] Step 8: Customizing the information provided

[1257] The server customizes the generated summary information based on the analyzed user's emotional state. It uses the generated summary and the user's emotional state as input. For example, if the user is "excited," it adds more appealing expressions to the summary. The output is the customized summary information.

[1258] Step 9: Provide a summary

[1259] The server sends customized summary information and associated advertising information to the device. The device uses the customized summary information and generated advertising information as input. The device displays this information to the user. The output is the customized summary information and advertising information provided to the user.

[1260] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1261] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1262] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1263] [Fourth Embodiment]

[1264] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1265] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1266] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1267] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1268] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1269] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1270] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1271] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1272] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1273] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1274] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1275] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1276] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1277] This invention relates to a system for collecting technical documents and intellectual property information, enabling users to efficiently obtain this information. The system is characterized by collecting information from a database, performing preprocessing, generating a summary using a generative artificial intelligence model, and providing it to the user.

[1278] System Programs and Processing

[1279] 1. Data Collection

[1280] The server connects to technical literature databases and intellectual property databases to collect relevant information.

[1281] The server queries the database using keywords such as "autonomous driving" and retrieves the corresponding data.

[1282] The server stores the retrieved data in a structured format such as JSON.

[1283] 2. Data preprocessing

[1284] The server converts the collected data into text format.

[1285] The server cleans up the data by removing noise from the text (e.g., HTML tags and special characters).

[1286] The server tokenizes the clean data and removes irrelevant stop words.

[1287] The server saves the pre-processed data to the database.

[1288] 3. Training the AI ​​model

[1289] The server uses the pre-processed dataset to train a generative artificial intelligence model.

[1290] The server uses generative artificial intelligence models (e.g., BERT or GPT) to learn important patterns in the data.

[1291] The server stores the trained model and deploys it for inference.

[1292] 4. Acceptance of user queries

[1293] The terminal receives search queries from the user through the input interface.

[1294] The user enters keywords such as "latest information on autonomous driving technology."

[1295] 5. Query analysis and extraction of corresponding data

[1296] The server receives the search query sent from the terminal and parses the keywords.

[1297] The server extracts relevant technical literature and intellectual property information from the database based on the query.

[1298] 6. Summary Generation

[1299] The server inputs the extracted data into a generative artificial intelligence model and generates a summary.

[1300] The server verifies the summary results to ensure that the main points are reflected.

[1301] 7. Providing a summary

[1302] The server converts the generated summary into JSON format and sends it to the terminal.

[1303] The device displays summary information to the user.

[1304] Specific example

[1305] Data Acquisition and Preprocessing

[1306] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to autonomous driving. The collected data is stored in JSON format, and after noise reduction and tokenization, it is stored in the database.

[1307] Model training and query analysis

[1308] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize important information related to autonomous driving technology. When a user enters "latest information on autonomous driving technology" as a query, the server analyzes this query and extracts relevant papers and patents.

[1309] Summary generation and delivery

[1310] The extracted information is summarized by a generative artificial intelligence model. For example, a summary might be generated stating, "The latest autonomous driving technologies include AI-powered real-time road surface detection and high-precision obstacle detection using laser sensors." The generated summary is sent from the server to the terminal, where the user can view it through the interface.

[1311] Based on the above, the present invention is a system that enables users to effectively collect and summarize technical documents and intellectual property information.

[1312] The following describes the processing flow.

[1313] Step 1:

[1314] The server connects to technical literature databases and intellectual property databases. For example, it uses APIs to collect data from the "technical literature database" and the "intellectual property database."

[1315] Step 2:

[1316] The server converts the collected data into JSON format and saves it to local or cloud storage. The collected data includes the title of the paper, authors, abstracts, patent numbers, and summaries of the inventions.

[1317] Step 3:

[1318] The server preprocesses the collected data. Specifically, it removes HTML tags and special characters from the text. Next, it tokenizes the cleansed data and removes unwanted words (stop words).

[1319] Step 4:

[1320] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model to begin the learning process.

[1321] Step 5:

[1322] The server evaluates the trained generative artificial intelligence model on a test dataset to verify its accuracy. Once the evaluation is complete, the model is deployed for inference.

[1323] Step 6:

[1324] The terminal receives search queries from users through its interface. For example, it might accept a query such as "latest trends in autonomous driving technology."

[1325] Step 7:

[1326] The server analyzes the search query received from the terminal and extracts keywords. For example, "autonomous driving technology" and "latest trends" might be extracted as keywords.

[1327] Step 8:

[1328] The server extracts relevant technical literature and intellectual property information from the database based on the analyzed query. The extracted data is temporarily loaded into memory.

[1329] Step 9:

[1330] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The summary includes key technical points, findings, and application examples.

[1331] Step 10:

[1332] The server converts the generated summary into JSON format and sends it to the terminal. Because the transmitted summary data is structured, it is displayed in a user-friendly format.

[1333] Step 11:

[1334] The terminal displays the received summary results in the user interface. Users can review the summary information corresponding to their search queries and quickly obtain the necessary technical information.

[1335] This series of processes allows users to efficiently collect the latest technical literature and intellectual property information, and obtain the necessary information in a summarized format.

[1336] (Example 1)

[1337] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1338] Because technical literature and intellectual property information are vast, it is difficult for individual users to efficiently collect the necessary information and summarize and understand the key points. Traditional systems primarily rely on manual information gathering and summarization, which is time-consuming and labor-intensive. Furthermore, the process of extracting and summarizing appropriate information based on search queries is often inefficient. Therefore, there is a need for a system that can automatically collect, preprocess, and summarize technical literature and intellectual property information, and efficiently provide it to users.

[1339] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1340] In this invention, the server includes means for collecting relevant information from a technical literature database and an intellectual property database; means for converting the collected information into text format, removing noise, and creating clean data; and means for tokenizing the clean data and preprocessing irrelevant information. This makes it possible to automatically collect and preprocess technical literature and intellectual property information to generate a dataset.

[1341] The server also includes means for training a generative artificial intelligence model using a preprocessed dataset, means for receiving search queries from users, means for analyzing the received search queries and extracting relevant information from technical literature databases and intellectual property databases, means for generating summaries of the extracted information using the generative artificial intelligence model, and means for converting the generated summaries into JSON format and providing them to users. This makes it possible to efficiently extract relevant information, summarize the key points, and provide them to users.

[1342] A "technical literature database" is a database that collects technical information, such as academic papers and technical reports, and stores it in a searchable format.

[1343] An "intellectual property database" is a database that collects and stores information related to intellectual property, such as patents, trademarks, and copyrights, in a searchable format.

[1344] A "generative artificial intelligence model" is an artificial intelligence model that has the ability to generate new data in tasks such as natural language processing and image generation.

[1345] "JSON format" is an abbreviation for JavaScript Object Notation, and it is a lightweight data exchange format for structuring and storing data.

[1346] "Noise reduction" is the process of removing unnecessary information (e.g., HTML tags and special characters) from data to generate clean data.

[1347] A "query" is a type of inquiry that a user enters in a database or search system to retrieve specific information.

[1348] "Tokenization" is the process of dividing text data into semantic units (tokens) such as words and phrases.

[1349] "Preprocessing" refers to the process of preparing data, such as cleaning up and standardizing its format, before analyzing the data or training models.

[1350] A "summary" is information that extracts the important points from a vast amount of information and presents them in a short format.

[1351] "Analysis" is the process of analyzing data and query content to understand and process their meaning and structure.

[1352] "Extraction" is the process of retrieving data or information based on specific conditions.

[1353] This invention relates to a system for efficiently collecting and processing technical literature and intellectual property information and providing it to users in a summarized format. The system collects relevant information from technical literature databases and intellectual property databases, performs preprocessing, generates summaries using a generative artificial intelligence model, and provides them to users.

[1354] Data collection

[1355] The server periodically accesses technical literature databases (e.g., IEEE Xplore) and intellectual property databases (e.g., USPTO) to collect relevant information. For example, the server automatically sends queries at 3:00 AM every day to retrieve the latest papers and patent information related to keywords such as "autonomous driving" and "AI technology." The collected information is temporarily stored in data storage in a structured JSON format.

[1356] Data preprocessing

[1357] The server converts the collected data into text format and removes noise. First, HTML tags and special characters are removed using regular expressions. Next, the text data is tokenized (divided into words and phrases), and irrelevant stop words (e.g., "a" and "the" in English) are removed using the NLTK library. The pre-processed data is then saved back to the database.

[1358] Training an AI model

[1359] The server trains a generative artificial intelligence (AI) model using a pre-processed dataset. For example, it trains a BERT model using the Hugging Face Transformers library. The server uses 80% of the data for training and 20% for validation. Once the training is complete, the model is stored in the Hugging Face model hub.

[1360] Acceptance of user queries

[1361] The terminal receives search queries from the user through an input interface. For example, it might be implemented as a web application running in a browser, where the user enters a query such as "latest autonomous driving technology" into the search bar. This query is sent to the server in real time.

[1362] Query analysis and extraction of corresponding data

[1363] The server receives search queries sent from the terminal and analyzes the keywords using TF-IDF (Term Frequency-Inverse Document Frequency). Based on the analysis results, the server extracts relevant information from the database. For example, it might pick out literature and patents related to "latest autonomous driving technology."

[1364] Summary generation

[1365] The server inputs the extracted literature and patent information into a generative artificial intelligence model to generate a summary. Specifically, it feeds the extracted text into a BERT model to generate a summarized text. During this process, the model undergoes multiple validations to ensure that all important points are included.

[1366] Summary

[1367] The server converts the generated summary into JSON format and sends it to the terminal. The terminal displays this summary information to the user. For example, in the user interface of a web application, the summary text is displayed in an easy-to-read format. The user can review the displayed summary and access detailed information as needed.

[1368] Specific example

[1369] Data Acquisition and Preprocessing

[1370] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to "autonomous driving." The collected data is stored in JSON format and, after preprocessing such as removing HTML tags and special characters and tokenizing the text, is stored in the database.

[1371] Model training and query analysis

[1372] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize key information related to "autonomous driving technology." When a user enters "latest information on autonomous driving technology" as a query, the server analyzes this query and extracts relevant papers and patents.

[1373] Summary generation and delivery

[1374] The extracted information is summarized by a generative artificial intelligence model. For example, a summary might be generated stating, "The latest autonomous driving technologies include AI-powered real-time road surface detection and high-precision obstacle detection using laser sensors." The generated summary is sent from the server to the terminal, where the user can view it in a web browser.

[1375] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1376] Step 1: Data collection from the database

[1377] The server connects to technical literature databases and intellectual property databases to collect information related to specific keywords (e.g., "autonomous driving" or "AI technology"). The input is a keyword query, and the output is the collected data in JSON format. Specifically, the server automatically sends queries at 3 AM every day and temporarily stores the retrieved papers and patent information in data storage.

[1378] Step 2: Data Preprocessing

[1379] The server converts the collected data into text format and removes noise. The input is collected data in JSON format, and the output is clean text data. Specifically, the server uses regular expressions to remove HTML tags and special characters, then tokenizes the text data into words and phrases, and removes irrelevant stop words using the NLTK library. This clean data is then stored again in the database.

[1380] Step 3: Training the AI ​​model

[1381] The server trains a generative artificial intelligence model using a preprocessed dataset. The input is a preprocessed dataset, and the output is a trained generative AI model. Specifically, the server trains a BERT model using the Hugging Face Transformers library, using 80% of the data for training and 20% for validation. The trained model is stored in the Hugging Face model hub.

[1382] Step 4: Accepting User Queries

[1383] The terminal receives search queries from the user through an input interface. The input is the user's search query (e.g., "latest autonomous driving technology"), and the output is that query being sent to the server. Specifically, the user enters the query into the search bar of the web application, and that query is sent to the server in real time.

[1384] Step 5: Query analysis and extraction of corresponding data

[1385] The server receives search queries sent from terminals and parses the keywords. The input is the user's search query, and the output is relevant information extracted from the database. Specifically, the server uses TF-IDF (Term Frequency-Inverse Document Frequency) to parse the query and extracts relevant technical documents and patent information from the database based on the results.

[1386] Step 6: Summary Generation

[1387] The server inputs the extracted information into a generative artificial intelligence model to generate a summary. The input includes extracted technical documents and patent information, and the output is the generated summary. Specifically, the server feeds the extracted text data into a BERT model to generate the summary. During this process, the model undergoes multiple validations to ensure that the key points are reflected.

[1388] Step 7: Provide a summary

[1389] The server converts the generated summary into JSON format and sends it to the terminal. The input is the generated summary text, and the output is the summary data sent to the terminal in JSON format. Specifically, the server converts the summary into JSON format and provides it to the user through the web application's user interface. The terminal displays the received summary to the user, who can access detailed information as needed.

[1390] (Application Example 1)

[1391] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1392] Traditionally, the collection and summarization of technical literature and intellectual property information was often done manually, requiring considerable time and effort. Furthermore, there was no way for users to instantly access the latest technical information about products within a virtual store, making effective information gathering and provision difficult. This resulted in reduced user convenience and impaired the efficiency of information gathering.

[1393] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1394] In this invention, the server includes means for collecting technical literature and intellectual property information from a database; means for preprocessing the collected data and removing irrelevant information from the text; means for training a generative artificial intelligence model using the preprocessed data; means for receiving search queries from users; means for extracting relevant information from the database based on the received search queries; means for generating a summary of the extracted information using the generative artificial intelligence model; means for providing the generated summary to the user; and means for enabling the user to instantly obtain the latest technical and intellectual property information about products in a virtual store using a smartphone or smart glasses. This makes it possible for users to efficiently collect technical literature and intellectual property information and instantly check the summarized information.

[1395] A "database" is a system for organizing, systematizing, and centrally managing information.

[1396] "Technical literature" refers to documents that describe research results and knowledge related to science and technology.

[1397] "Intellectual property information" refers to information concerning intellectual property rights such as patents, trademarks, and copyrights.

[1398] "Preprocessing" is the process of removing noise and irrelevant information from raw data and converting it into an analyzable format.

[1399] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates text and other data based on training data.

[1400] A "search query" is a keyword or phrase that a user uses to search for specific information.

[1401] A "summary" is a short, concise version of a longer text, extracting the most important information from it.

[1402] A "smartphone" is a portable device that, in addition to telephone functionality, also possesses advanced computer capabilities.

[1403] "Smart glasses" are glasses-type devices that display information in the field of view and have the ability to connect to the internet.

[1404] A "virtual store" is an online shop that exists on the internet, where users can browse and purchase products in a virtual space.

[1405] A "product" is an item or service that is manufactured and supplied to consumers.

[1406] This invention relates to a system that collects technical literature and intellectual property information and provides summarized information using a generative artificial intelligence model. This system enables users, particularly in virtual stores, to instantly obtain the latest technical and intellectual property information about products using a smartphone or smart glasses.

[1407] System Overview

[1408] 1. Collection from databases

[1409] The server connects to technical literature databases and intellectual property databases to collect relevant information. For example, it queries the databases using keywords such as "autonomous driving" and retrieves corresponding data. The retrieved data is stored in a structured format such as JSON.

[1410] 2. Data preprocessing

[1411] The server converts the collected data into text format. It then cleans the data by removing noise (e.g., HTML tags and special characters). Next, it tokenizes the text data and removes irrelevant stop words. The pre-processed data is then stored back in the database.

[1412] 3. Training the AI ​​model

[1413] The server uses preprocessed data to train a generative artificial intelligence model. This model may be one such model, such as BERT or GPT. Once trained, the model learns important patterns and is deployed for inference.

[1414] 4. Acceptance of user queries

[1415] The terminal receives search queries from users through an input interface. For example, a user might enter keywords such as "latest information on autonomous driving technology."

[1416] 5. Query analysis and extraction of corresponding data

[1417] The server receives search queries sent from the terminal and analyzes the keywords. Next, it extracts relevant technical literature and intellectual property information from the database based on the query.

[1418] 6. Summary Generation

[1419] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The generated summary is then verified to ensure that the main points are reflected before being provided to the user.

[1420] 7. Providing a summary

[1421] The server sends the summarized results, converted to JSON format, to the terminal. The terminal then displays the generated summarized information to the user.

[1422] Specific example

[1423] For example, if the user wants to obtain information about "autonomous driving," the server sends queries to technical literature databases and intellectual property databases, collects papers and patents related to autonomous driving, and generates clean data. A trained generative artificial intelligence model generates a summary from this data, such as "the latest autonomous driving technologies include AI-based real-time road surface detection and high-precision obstacle detection using laser sensors." Users can instantly check this summary information via their smartphone or smart glasses. This invention enables rapid information provision within virtual stores.

[1424] Example of a prompt

[1425] The following prompts are used as input to the generative artificial intelligence model:

[1426] Summarize the latest information on autonomous driving technology.

[1427] The above describes the embodiments of the present invention. This system allows users to efficiently collect and verify the latest technical and intellectual property information within a virtual store.

[1428] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1429] Step 1:

[1430] The server connects to technical literature databases and intellectual property databases to collect relevant information. The server queries the databases using keywords (e.g., "autonomous driving") and stores the retrieved data in JSON format.

[1431] Input: Keyword query

[1432] Output: Technical document data and intellectual property information in JSON format

[1433] Specific operation: The server accesses the database via an API and retrieves and saves the relevant data.

[1434] Step 2:

[1435] The server preprocesses the collected data. Specifically, it converts it to text data, removes noise (e.g., HTML tags and special characters), and removes irrelevant stop words. It also performs tokenization. The preprocessed data is then stored again in the database.

[1436] Input: Technical document data and intellectual property information in JSON format

[1437] Output: Preprocessed text data

[1438] Specific operation: The server performs text data cleanup, tokenization, and stop word removal.

[1439] Step 3:

[1440] The server trains generative artificial intelligence models using preprocessed data. Specifically, it inputs data into models such as BERT and GPT to learn patterns. Once the training is complete, the models are stored on the server and deployed for inference.

[1441] Input: Preprocessed text data

[1442] Output: Trained generative artificial intelligence model

[1443] Specific operation: The server inputs data into the model and executes and completes the training process.

[1444] Step 4:

[1445] The terminal receives search queries from the user. The user enters search keywords (e.g., "latest information on autonomous driving technology") through the interface.

[1446] Input: Search keyword (user query)

[1447] Output: Search query (sent from terminal to server)

[1448] Specific operation: The terminal receives queries through the user interface and sends them to the server.

[1449] Step 5:

[1450] The server receives search queries sent from terminals and analyzes the keywords. Based on the queries, the server extracts relevant technical literature and intellectual property information from its database.

[1451] Input: Search query from the user

[1452] Output: Extracted technical literature and intellectual property information

[1453] Specific operation: The server parses the query and quickly extracts relevant information from the database.

[1454] Step 6:

[1455] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The generated summary is then verified to ensure that the main points are reflected before being provided to the user.

[1456] Input: Extracted technical literature and intellectual property information

[1457] Output: Summarized information

[1458] Specific operation: The server uses a trained generative artificial intelligence model to generate and verify a summary of the extracted data.

[1459] Step 7:

[1460] The server converts the generated summary into JSON format and sends it to the terminal. The terminal then displays the summary information to the user.

[1461] Input: Summary information (generated by a generative AI model)

[1462] Output: Summary information in JSON format (sent from server to terminal)

[1463] Specific operation: The server encodes the summary information and sends it to the terminal. The terminal displays the summary information to the user.

[1464] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1465] This invention relates to a system for collecting technical literature and intellectual property information, enabling users to efficiently obtain this information. The system is characterized by collecting information from a database, performing preprocessing, generating summaries using a generative artificial intelligence model, and providing them to the user. Furthermore, it incorporates an emotion engine that analyzes the user's emotions, and adjusts and customizes the information based on the user's emotional state.

[1466] System Programs and Processing

[1467] 1. Data Collection

[1468] The server connects to technical literature databases and intellectual property databases to collect relevant information. For example, it uses APIs to retrieve "technical literature" and "intellectual property information" from the databases.

[1469] 2. Data preprocessing

[1470] The server converts the collected data into text format and removes noise (e.g., HTML tags and special characters). Next, it tokenizes the clean data and removes irrelevant stop words.

[1471] 3. Training the AI ​​model

[1472] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model, and the learning process proceeds.

[1473] 4. Acceptance of user queries

[1474] The terminal receives search queries from the user through its interface. For example, the user might enter the query, "Latest trends in autonomous driving technology."

[1475] 5. Query analysis and extraction of corresponding data

[1476] The server analyzes the search query received from the terminal and extracts keywords. Next, it extracts relevant technical literature and intellectual property information from the database based on the query.

[1477] 6. Summary Generation

[1478] The server inputs the extracted data into a generative artificial intelligence model to generate a summary. The summary includes key technical points, findings, and application examples.

[1479] 7. Emotion analysis

[1480] The device collects emotional data from the user's interface operations and input.

[1481] The server uses an emotion engine to analyze the user's emotions. For example, it uses text analysis to determine whether the user is experiencing emotions such as "anxiety" or "interest."

[1482] 8. Customization of information provision

[1483] The server customizes the generated summary information based on the analyzed emotional state. For example, if the user is feeling "anxious," it adds reassuring phrases.

[1484] 9. Providing a summary

[1485] The server sends a customized summary to the terminal.

[1486] The device displays summary information to the user.

[1487] Specific example

[1488] Data Acquisition and Preprocessing

[1489] For example, when acquiring information on "autonomous driving," the server sends queries to a technical literature database and an intellectual property database to collect papers and patents related to autonomous driving. The collected data is stored in JSON format, and after noise reduction and tokenization, it is stored in the database.

[1490] Model training and query analysis

[1491] The generative artificial intelligence model is trained using pre-processed data. The trained model is designed to efficiently summarize important information related to autonomous driving technology. When a user enters "latest information on autonomous driving technology" as a query, the server analyzes this query and extracts relevant papers and patents.

[1492] Summarization generation and sentiment analysis

[1493] The extracted information is summarized by a generative artificial intelligence model. Simultaneously, the terminal collects the user's emotions when they enter queries, and the server analyzes those emotions using an emotion engine.

[1494] Customization and delivery

[1495] The summarized information is customized based on the user's emotional state. For example, if the user indicates "anxiety," the summary will include expressions that emphasize stability and reliability. Finally, the customized summary information is sent from the server to the terminal and displayed to the user.

[1496] In summary, the present invention is a system that enables users to effectively collect and summarize technical literature and intellectual property information, as well as to provide customized information based on the user's emotional state.

[1497] The following describes the processing flow.

[1498] Step 1:

[1499] The server connects to technical literature databases and intellectual property databases, and uses APIs and queries to collect "technical literature" and "intellectual property information." For example, it retrieves data related to "autonomous driving."

[1500] Step 2:

[1501] The server converts the collected data into JSON format and stores it securely in local or cloud storage.

[1502] Step 3:

[1503] The server preprocesses the collected data. Specifically, it removes HTML tags and special characters from the text. Furthermore, it performs natural language processing such as tokenization and stop word removal to clean the data.

[1504] Step 4:

[1505] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using a pre-processed dataset. The training dataset is then fed into the model to learn important patterns and features.

[1506] Step 5:

[1507] The server evaluates the trained generative artificial intelligence model on a test dataset to verify its accuracy. Once the evaluation is complete, the model is deployed for production use.

[1508] Step 6:

[1509] The terminal receives search queries from the user through its interface. For example, the user enters the query "latest trends in autonomous driving technology."

[1510] Step 7:

[1511] The terminal collects emotional data from the user's facial expressions, typing speed, and context when they enter queries. The collected emotional data is sent to the server.

[1512] Step 8:

[1513] The server analyzes the search queries received from the terminal and extracts "autonomous driving technology" and "latest trends" as keywords. Next, it extracts technical literature and intellectual property information related to these keywords from the database.

[1514] Step 9:

[1515] The server inputs the extracted information into a generative artificial intelligence model to generate a summary. The summary includes key technical points, findings, and application examples.

[1516] Step 10:

[1517] The server uses an emotion engine to analyze the transmitted emotion data. For example, it might identify the user's emotion as "anxiety."

[1518] Step 11:

[1519] The server customizes the summary generated based on the sentiment analysis results. For example, if "anxiety" is identified, it adds reassuring wording to the summary.

[1520] Step 12:

[1521] The server converts the customized summary into JSON format and sends it to the terminal.

[1522] Step 13:

[1523] The device displays the received, customized summary results in the user interface. Users can review the summary information for their queries and obtain information tailored to their emotional state.

[1524] This series of processes allows users to efficiently collect the latest technical literature and intellectual property information and confidently utilize the summarized information. The system also takes the user's emotional state into account, enabling more personalized information delivery.

[1525] (Example 2)

[1526] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1527] In conventional technical information gathering systems, it was difficult to efficiently collect relevant information from vast databases and summarize that information. Furthermore, the summarized information did not adapt to the user's emotional state, resulting in a decline in the quality of information provided. As a result, users were unable to obtain the necessary information quickly and appropriately, leading to wasted effort and time.

[1528] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1529] In this invention, the server includes means for collecting technical documents and intellectual property information from a database, means for preprocessing the collected data and removing irrelevant information from the text, and means for training a generative artificial intelligence model using the preprocessed data. This makes it possible to efficiently preprocess the collected data and train the generative artificial intelligence model with high accuracy. Furthermore, by including means for extracting relevant information from the database based on a received search query and generating a summary of that information using the generative artificial intelligence model, means for analyzing the user's emotional state during summary generation and customizing the summary information based on that state, and means for providing the generated summary to the user, it becomes possible to provide high-quality information adapted to the user's emotions, thereby improving the efficiency and accuracy of information collection.

[1530] A "database" is an information system that systematically stores multiple pieces of data, making it possible to search for and retrieve them.

[1531] "Technical literature" refers to documents such as academic papers, research reports, and technical reports related to a specific technical field.

[1532] "Intellectual property information" refers to data and documents related to intellectual property rights, such as patents, trademarks, and copyrights.

[1533] "Preprocessing" refers to a series of operations that transform data into a format suitable for analysis and model training.

[1534] A "generative artificial intelligence model" is an artificial intelligence model that can generate new text or information based on input data.

[1535] "Tokenization" is the process of dividing text data into meaningful units such as words or phrases.

[1536] A "stop word" is a common word that is considered low in importance and is excluded in text analysis and natural language processing.

[1537] A "search query" is a keyword or phrase that a user enters to search for information.

[1538] A "summary" is a shortened version of a document or piece of information, containing the main points extracted from it.

[1539] "Sentiment analysis" is the process of determining a user's emotional state from their text and behavior.

[1540] "Customization" refers to modifying or adjusting information and services to suit the specific needs and circumstances of the user.

[1541] "User interface" is a general term for the screens and operating methods that users use to interact with a system.

[1542] "Noise" refers to data or information that is unnecessary or harmful to the analysis.

[1543] Hugging Face is an organization that provides open-source libraries and tools for natural language processing.

[1544] The "Transformers library" is an open-source software library used for building and training generative artificial intelligence models.

[1545] Elasticsearch is a full-text search engine and analytics engine, a system designed for high-speed searching of large amounts of data.

[1546] "spaCy" is a high-performance natural language processing (NLP) library used for text analysis and model building.

[1547] "BeautifulSoup" is a Python library for parsing HTML and XML files and extracting specific information.

[1548] TextBlob is a Python library for easily performing sentiment analysis and classification of text data.

[1549] "VADER" is a Python library specifically designed for text sentiment analysis.

[1550] This invention relates to a system for efficiently collecting, summarizing, and providing technical literature and intellectual property information to users. This system is characterized by collecting information from a database, performing preprocessing, generating summaries using a generative artificial intelligence model, and providing information customized based on the user's sentiment.

[1551] Data collection

[1552] The server connects to technical literature databases and intellectual property databases via APIs to collect relevant information. For example, it collects papers and patent data related to "autonomous driving" from databases such as IEEE Xplore and Google Patents. The server authenticates using an API key, sends search queries, and retrieves data.

[1553] Specific example: The server sends "self-driving cars" as a search query to the IEEE Xplore API to retrieve a list of papers related to autonomous driving.

[1554] Data preprocessing

[1555] The server converts the collected data into text format and removes noise such as HTML tags and special characters using Python's BeautifulSoup library. Next, the clean text data is tokenized, and irrelevant stop words are removed using Python's NLTK library.

[1556] Specific example: Perform the operation of removing HTML tags from acquired research paper data and extracting only the text portion. In the text "The advancements in self-driving cars include sensors and AI technology.", stop words such as "The" and "in" are removed, and the tokens "advancements, self-driving, cars, include, sensors, AI, technology" are extracted.

[1557] Training an AI model

[1558] The server uses preprocessed data to train generative artificial intelligence models (e.g., BERT or GPT). The Hugging Face Transformers library is used to train the generative AI models. The GPT-3 model is trained using preprocessed text data on autonomous driving technology, aiming for highly accurate summary generation.

[1559] Acceptance of user queries

[1560] The terminal accepts search queries from users through a GUI (Graphical User Interface). For example, a user might enter the query "latest trends in autonomous driving technology" and click the search button.

[1561] Specific example: The user enters "latest trends in autonomous driving technology" and presses the search button.

[1562] Query analysis and extraction of corresponding data

[1563] The server parses the search query received from the terminal and extracts keywords from the query using the Python spaCy library. Next, it searches the database using Elasticsearch and extracts relevant technical documents and patent information.

[1564] Specific example: Extract the keyword "autonomous driving technology" from the query "latest trends in autonomous driving technology," and use Elasticsearch to search for papers and patent information related to this keyword.

[1565] Summary generation

[1566] The server inputs the extracted data into a generative artificial intelligence model to generate a summary that includes key technical points, discoveries, and application examples. For example, it can use a GPT-3 model to quickly summarize information and provide it to the user.

[1567] Specific example: Input extracted technical documents and patent information into GPT-3 to generate a "Summary of the Latest Trends in Autonomous Driving Technology."

[1568] Emotion analysis

[1569] The device collects emotional data from the user's text input and actions. For example, it uses an emotional analysis library (e.g., TextBlob or VADER) to analyze whether the user is experiencing emotions such as "anxiety" or "interest."

[1570] Specific example: If a user enters a keyword like "difficult," the TextBlob is used to analyze the emotion "anxiety."

[1571] Customization of information provision

[1572] The server customizes the summary information generated based on the analyzed emotional state. For example, if the user is feeling "anxious," it adds reassuring phrases and detailed information.

[1573] Specific example: Add a sentence to the summary such as, "This technology has already been safely implemented by major manufacturers and has a proven track record."

[1574] Summary

[1575] The server sends customized summary information to the terminal, which then displays it to the user.

[1576] Specific example: Display a customized summary in the user interface of a browser.

[1577] Through the steps described above, this system can efficiently collect and analyze technical literature and intellectual property information, and provide information tailored to the user's emotions. This allows users to quickly obtain the necessary information and deepen their understanding.

[1578] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1579] Program processing flow

[1580] Step 1: Data Collection

[1581] Input: Search queries to technical literature databases and intellectual property databases

[1582] The server connects to technical literature databases and intellectual property databases via APIs and collects data related to the information requested by the user. In this process, the server authenticates using an API key and sends search queries to retrieve data. For example, it might send "self-driving cars" as a query to the IEEE Xplore API.

[1583] Output: Collected technical literature data and intellectual property data

[1584] Step 2: Data Preprocessing

[1585] Input: Collected technical literature data and intellectual property data

[1586] The server converts the collected data into text format and removes noise such as HTML tags and special characters. The Python BeautifulSoup library is used for this process. Furthermore, the clean text data is tokenized, and irrelevant stop words are removed using the NLTK library. In this way, clean data suitable for analysis is obtained.

[1587] Output: Pre-processed clean data

[1588] Step 3: Training the AI ​​model

[1589] Input: Pre-processed clean data

[1590] The server uses preprocessed data to train generative artificial intelligence models (e.g., BERT or GPT). This training utilizes the Hugging Face Transformers library. The generative AI models learn to generate highly accurate summaries based on a large amount of technical literature data.

[1591] Output: Trained generative AI model

[1592] Step 4: Accepting User Queries

[1593] Input: User's search query

[1594] The terminal accepts search queries from users via a GUI. The user enters a query (e.g., "latest trends in autonomous driving technology") into the text input field and clicks the search button.

[1595] Output: User's search query

[1596] Step 5: Query analysis and extraction of corresponding data

[1597] Input: User's search query

[1598] The server parses the search query received from the terminal and extracts keywords from the query using the Python spaCy library. Next, it searches the database using Elasticsearch to extract relevant technical documents and patent information. For example, it extracts the keyword "autonomous driving technology" from the query "latest trends in autonomous driving technology" and performs a search.

[1599] Output: Extracted technical literature and patent information

[1600] Step 6: Summary Generation

[1601] Input: Extracted technical documents and patent information

[1602] The server inputs the extracted data into a generative artificial intelligence model to generate a summary that includes key technical points, discoveries, and application examples. For example, it uses a GPT-3 model to quickly summarize information and provide it to the user.

[1603] Output: Summarized information

[1604] Step 7: Emotion Analysis

[1605] Input: User query input and operation information

[1606] The device collects emotional data from the user's text input and actions. For example, it uses an emotional analysis library (such as TextBlob or VADER) to determine if the user is experiencing emotions like "anxiety" or "interest." If the user types "difficult," the device analyzes that emotion.

[1607] Output: User's emotional state

[1608] Step 8: Customizing the information provided

[1609] Input: Summarized information and user's emotional state

[1610] The server customizes the summary information generated based on the analyzed emotional state. If the user is feeling "anxious," reassuring phrases and detailed information are added to the summary. For example, a sentence like, "This technology has already been securely implemented by major manufacturers and has a proven track record," might be added to the summary.

[1611] Output: Customized summary information

[1612] Step 9: Provide a summary

[1613] Input: Customized summary information

[1614] The server sends customized summary information to the terminal. The terminal displays the received summary information in its user interface, providing the information to the user.

[1615] Output: Summary information displayed to the user

[1616] Through the steps outlined above, this system enables the effective collection, analysis, and summarization of technical literature and intellectual property information, as well as the provision of information based on user sentiment. This allows users to quickly obtain the necessary information and make appropriate decisions.

[1617] (Application Example 2)

[1618] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1619] In today's world, users are required to acquire advanced information quickly and efficiently. However, a vast amount of technical literature and intellectual property information exists, and extracting and summarizing highly relevant information from it requires considerable time and effort. Furthermore, the information provided may lack effectiveness for the recipient because it is not appropriately customized according to the user's emotions and circumstances. In addition, advertising also faces the challenge of not providing appropriate information that is tailored to the user's situation.

[1620] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting technical literature and intellectual property information from a database, means for preprocessing the collected data and removing irrelevant information from the text, means for training a generative artificial intelligence model using the preprocessed data, means for customizing a summary generated based on the user's emotional state, and means for providing the user with advertising information including the customized summary. As a result, the user can efficiently obtain optimized technical literature and intellectual property information, and the receptivity of the information is improved because the most appropriate advertising information is provided according to the user's emotions and situation.

[1621] A "database" is a collection of information, including technical documents and intellectual property information. The information is stored in a searchable format.

[1622] "Technical documents" are materials that describe research results and technical explanations related to science and technology. This includes academic papers and technical reports.

[1623] "Intellectual property information" refers to information concerning intellectual property rights such as patents, trademarks, and copyrights. This includes patent application documents and registration certificates.

[1624] "Preprocessing" is the process of removing irrelevant information and noise from collected data and converting the data into an analyzable format. This also includes the process of removing HTML tags and special characters from text.

[1625] A "generative artificial intelligence model" is a machine learning model that has the ability to generate new information based on large amounts of data. Examples include BERT and GPT.

[1626] A "search query" is a request for input from a user to obtain specific information. It consists of keywords or questions.

[1627] "Summary generation" is the process of shortening long texts or data and extracting the main points. This is done automatically using artificial intelligence models.

[1628] "Emotional state" refers to the psychological state a user is in when receiving information. Specifically, it includes "excitement," "anxiety," and "interest."

[1629] "Customization" refers to adjusting and modifying information to suit the specific needs and feelings of the user.

[1630] "Advertising information" refers to information used to promote products and services. It is provided in an optimized format based on the user's interests and emotional state.

[1631] This invention relates to a system that collects technical literature and intellectual property information from a database and provides information optimized for the user. This system preprocesses the collected data, generates summaries using a generative artificial intelligence model, further customizes the information based on the user's emotional state, and provides advertising information along with it.

[1632] Hardware and software usage

[1633] The system is implemented using the following hardware and software:

[1634] Server: Responsible for large-scale data collection, preprocessing, training of generative artificial intelligence models, and summary generation.

[1635] Terminal: Responsible for receiving search queries from users and collecting sentiment data.

[1636] Generative artificial intelligence models: For example, BERT or GPT are used.

[1637] Emotion analysis engine: Emotion analysis is performed using the Hugging Face Transformers library.

[1638] Database: Stores technical literature and intellectual property information, and provides data via API.

[1639] Data processing and data calculation

[1640] Data Acquisition and Preprocessing

[1641] The server collects technical literature and intellectual property information from a database via an API. The collected data is then stored in JavaScript Object Notation (JSON) format and converted to text format. The text undergoes preprocessing such as noise reduction, tokenization, and stop word removal.

[1642] Training of generative artificial intelligence models

[1643] Using preprocessed data, the server trains a generative artificial intelligence model. This training is performed to generate summaries related to a specific technical domain. For example, PyTorch is used to train the model.

[1644] Emotion analysis

[1645] The device collects emotional data in real time as the user enters search queries. This emotional data is input into an emotional analysis engine through text analysis. The emotional analysis engine detects what emotional state the user is in, such as "excitement," "interest," or "anxiety."

[1646] Customized information provision

[1647] The server customizes the generated summary information based on sentiment analysis results. For example, it adds reassuring or attention-grabbing phrases to the summary. It then sends the customized summary along with advertising information to the device.

[1648] Specific example

[1649] When a user enters the query "latest smartphone technology," the server collects relevant information from technical literature and intellectual property databases. The collected data is then preprocessed and summarized by a generative artificial intelligence model. After the summary is generated, an emotion analysis engine analyzes the user's emotional state, and the summary information is customized based on the results.

[1650] For example, if a user expresses excitement, the summary information will include phrases that emphasize the appeal and future potential of the new technology. Finally, the customized summary information and related advertisements are sent to the device and displayed to the user.

[1651] Example of a prompt:

[1652] "Could you tell me about the latest trends in smartphone technology?"

[1653] The above describes the specific forms for carrying out the invention.

[1654] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1655] Step 1: Data Collection

[1656] The server collects relevant information from technical literature databases and intellectual property information databases via APIs. The input is a query containing keywords related to the technical literature and intellectual property information to be collected. The server stores the data collected via the API in JSON format. The output is a dataset of the collected technical literature and intellectual property information.

[1657] Step 2: Data Preprocessing

[1658] The server preprocesses the data collected in Step 1. Specifically, it removes noise such as HTML tags and special characters from the text and performs tokenization. The collected dataset of technical documents and intellectual property information is used as input. The output is clean data that has been de-noised and tokenized.

[1659] Step 3: Training the Generative AI Model

[1660] The server trains a generative artificial intelligence model (e.g., BERT or GPT) using preprocessed data. The input is a preprocessed dataset. Training is performed using machine learning libraries such as PyTorch. The output is a generative artificial intelligence model trained to generate summaries corresponding to a specific technology domain.

[1661] Step 4: Accepting User Queries

[1662] The terminal receives search queries from the user through its interface. The input is the search query entered by the user (e.g., "Please tell me about the latest trends in smartphone technology."). The output is the received search query.

[1663] Step 5: Query analysis and data extraction

[1664] The server parses the search query received from the terminal and extracts keywords. It uses the user's search query as input. Next, the server extracts relevant technical literature and intellectual property information from the database based on the query. The output is a set of relevant technical literature and intellectual property information.

[1665] Step 6: Summary Generation

[1666] The server inputs extracted technical literature and intellectual property information into a generative artificial intelligence model to generate a summary. The extracted dataset and the trained generative AI model are used as input. The output is a summary containing the main points.

[1667] Step 7: Emotion Analysis

[1668] The terminal collects sentiment data when the user enters a search query. The input consists of user interface interactions and input text. The server inputs the collected sentiment data into a sentiment analysis engine to analyze the user's emotional state (e.g., "excited," "anxious"). The output is the user's emotional state.

[1669] Step 8: Customizing the information provided

[1670] The server customizes the generated summary information based on the analyzed user's emotional state. It uses the generated summary and the user's emotional state as input. For example, if the user is "excited," it adds more appealing expressions to the summary. The output is the customized summary information.

[1671] Step 9: Provide a summary

[1672] The server sends customized summary information and associated advertising information to the device. The device uses the customized summary information and generated advertising information as input. The device displays this information to the user. The output is the customized summary information and advertising information provided to the user.

[1673] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1674] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1675] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1676] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1677] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1678] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1679] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1680] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1681] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1682] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1683] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1684] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1685] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1686] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1687] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1688] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1689] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1690] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1691] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1692] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1693] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1694] The following is further disclosed regarding the embodiments described above.

[1695] (Claim 1)

[1696] Means for collecting technical literature and intellectual property information from a database,

[1697] A means of preprocessing the collected data and removing irrelevant information from the text,

[1698] A means for training a generative artificial intelligence model using preprocessed data,

[1699] A means of receiving search queries from users,

[1700] A means for extracting relevant information from a database based on a received search query,

[1701] A means for generating a summary of extracted information using a generative artificial intelligence model,

[1702] A means of providing the generated summary to the user,

[1703] A system that includes this.

[1704] (Claim 2)

[1705] The system according to claim 1, which uses keyword queries when collecting technical documents and intellectual property information from a database.

[1706] (Claim 3)

[1707] The system according to claim 1, wherein noise reduction and tokenization are performed in the preprocessing process.

[1708] "Example 1"

[1709] (Claim 1)

[1710] Means for collecting relevant information from technical literature databases and intellectual property databases,

[1711] A method for converting collected information into text format, removing noise, and creating clean data,

[1712] A preprocessing method that tokenizes clean data and removes irrelevant information,

[1713] A means for training a generative artificial intelligence model using a preprocessed dataset,

[1714] A means of receiving search queries from users,

[1715] A means for analyzing received search queries and extracting relevant information from technical literature databases and intellectual property databases,

[1716] A means for generating a summary of extracted information using a generative artificial intelligence model,

[1717] A means of converting the generated summary into JSON format and providing it to the user,

[1718] A system that includes this.

[1719] (Claim 2)

[1720] The system according to claim 1, which uses keyword queries when collecting information from a technical literature database and an intellectual property database.

[1721] (Claim 3)

[1722] The system according to claim 1, which performs noise reduction and tokenization of HTML tags and special characters in a preprocessing process.

[1723] "Application Example 1"

[1724] (Claim 1)

[1725] Means for collecting technical literature and intellectual property information from a database,

[1726] A means of preprocessing the collected data and removing irrelevant information from the text,

[1727] A means for training a generative artificial intelligence model using preprocessed data,

[1728] A means of receiving search queries from users,

[1729] A means for extracting relevant information from a database based on a received search query,

[1730] A means for generating a summary of extracted information using a generative artificial intelligence model,

[1731] A means of providing the generated summary to the user,

[1732] A means by which users can instantly obtain the latest technical and intellectual property information about products within a virtual store using a smartphone or smart glasses,

[1733] A system that includes this.

[1734] (Claim 2)

[1735] The system according to claim 1, which uses keyword queries when collecting technical documents and intellectual property information from a database.

[1736] (Claim 3)

[1737] The system according to claim 1, wherein noise reduction and tokenization are performed in the preprocessing process.

[1738] "Example 2 of combining an emotion engine"

[1739] (Claim 1)

[1740] Means for collecting technical literature and intellectual property information from a database,

[1741] A means of preprocessing the collected data and removing irrelevant information from the text,

[1742] A means for training a generative artificial intelligence model using preprocessed data,

[1743] A means of receiving search queries from users,

[1744] A means for extracting relevant information from a database based on a received search query,

[1745] A means for generating a summary of extracted information using a generative artificial intelligence model,

[1746] A means for analyzing the user's emotional state during summary generation and customizing the summary information based on that state,

[1747] A means of providing the generated summary to the user,

[1748] A system that includes this.

[1749] (Claim 2)

[1750] The system according to claim 1, which uses keyword queries when collecting technical documents and intellectual property information from a database.

[1751] (Claim 3)

[1752] The system according to claim 1, wherein noise reduction and tokenization are performed in the preprocessing process.

[1753] "Application example 2 when combining with an emotional engine"

[1754] (Claim 1)

[1755] Means for collecting technical literature and intellectual property information...

Claims

1. Means for collecting technical literature and intellectual property information from a database, A means of preprocessing the collected data and removing irrelevant information from the text, A means for training a generative artificial intelligence model using preprocessed data, A means of receiving search queries from users, A means for extracting relevant information from a database based on a received search query, A means for generating a summary of extracted information using a generative artificial intelligence model, A means of providing the generated summary to the user, A system that includes this.

2. The system according to claim 1, which uses keyword queries when collecting technical documents and intellectual property information from a database.

3. The system according to claim 1, wherein noise reduction and tokenization are performed in the preprocessing process.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A